Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent–Child Interaction
Honorable MentionAuthors
Paper Title
Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent–Child Interaction
Publication Info
- Topic area: Alignment of multimodal large language models (MLLMs) with expert reasoning for analyzing parent–child interactions.
- Keywords: MLLMs, joint attention, parent–child interaction, speech-language pathologists, multimodal AI, human-centered AI, expert alignment, gaze, vocalization, action.
Background and Problem
- Problem / challenge: Current AI systems lack the ability to emulate the nuanced observational and judgment practices of speech-language pathologists (SLPs) in analyzing complex social behaviors like joint attention in parent–child interactions.
- Significance: Addressing this gap could improve access to developmental assessments, particularly for families in underserved communities, and enhance AI’s role in supporting early childhood development.
- Motivation and related work: Prior AI systems have focused on therapy material creation and caregiver engagement but have not aligned with SLPs’ evaluative expertise. Studies show that MLLMs can interpret multimodal behavior but struggle with socially nuanced tasks. This paper investigates whether MLLMs can align with SLPs’ reasoning to assess joint attention.
Solution
- Proposed approach: A two-stage MLLM-based system that separates observation (behavioral cues like gaze, action, vocalization) from judgment (interaction quality assessment).
- Novelty:
- Characterization of SLPs’ reasoning processes for joint attention using interviews and video annotations.
- Development of an MLLM system achieving up to 85% accuracy in observation and 64% accuracy in judgment compared to expert labels.
- Creation of a dataset with expert-labeled joint attention behaviors and design guidelines for aligning MLLMs with SLP workflows.
- Procedure and key techniques:
- Conduct interviews and annotation sessions with three SLPs to identify key behavioral dimensions (gaze, action, vocalization).
- Translate expert heuristics into structured prompts for MLLMs to extract behavioral segments from videos.
- Evaluate MLLM judgment accuracy using zero-shot and many-shot prompting strategies with GPT-4.1.
Results
- Concrete findings:
- Observation-level alignment: MLLMs achieved 81% accuracy in identifying gaze, action, and vocalization cues.
- Judgment-level alignment: Many-shot prompting improved accuracy to 64%, macro-F1 to 0.44, and Cohen’s κ to 0.26.
- Advantage over baselines: Structured prompts improved behavioral description accuracy from 62% to 81%, and judgment alignment outperformed zero-shot prompting.
- Experiments / evaluation:
- Dataset: 25 YouTube videos of parent–child interactions annotated by three SLPs.
- Metrics: Accuracy, macro-precision, macro-recall, macro-F1, Cohen’s κ.
- Models: Gemini 2.5 Pro for observation, GPT-4.1 for judgment.
- Limitations and future work:
- Small sample size of three SLPs limits generalizability.
- Dataset skewed toward neurotypical interactions, with few examples of poor joint attention.
- Stage 2 results based on manually corrected descriptions, representing an upper bound on performance.
- Future work should expand datasets to include neurodiverse populations and larger expert cohorts.
Summary
This study explores the alignment of multimodal large language models (MLLMs) with speech-language pathologists (SLPs) in analyzing joint attention during parent–child interactions. The proposed two-stage system separates observation (behavioral cues like gaze, action, vocalization) from judgment (interaction quality assessment). Results show robust alignment at the observation level (81% accuracy) but modest alignment at the judgment level (64% accuracy with many-shot prompting). Findings highlight the feasibility of using MLLMs to support expert workflows and suggest design opportunities for human-centered AI systems that augment professional expertise in developmental assessments.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)