Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent–Child Interaction

Honorable Mention
Human-LLM CollaborationHuman Pose & Activity RecognitionChild-Computer Interaction DesignSpeech-Language Pathologists & AudiologistsUniversity Professors & Researchers

Paper Title

Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent–Child Interaction

Publication Info

  • Topic area: Alignment of multimodal large language models (MLLMs) with expert reasoning for analyzing parent–child interactions.
  • Keywords: MLLMs, joint attention, parent–child interaction, speech-language pathologists, multimodal AI, human-centered AI, expert alignment, gaze, vocalization, action.

Background and Problem

  • Problem / challenge: Current AI systems lack the ability to emulate the nuanced observational and judgment practices of speech-language pathologists (SLPs) in analyzing complex social behaviors like joint attention in parent–child interactions.
  • Significance: Addressing this gap could improve access to developmental assessments, particularly for families in underserved communities, and enhance AI’s role in supporting early childhood development.
  • Motivation and related work: Prior AI systems have focused on therapy material creation and caregiver engagement but have not aligned with SLPs’ evaluative expertise. Studies show that MLLMs can interpret multimodal behavior but struggle with socially nuanced tasks. This paper investigates whether MLLMs can align with SLPs’ reasoning to assess joint attention.

Solution

  • Proposed approach: A two-stage MLLM-based system that separates observation (behavioral cues like gaze, action, vocalization) from judgment (interaction quality assessment).
  • Novelty:
    1. Characterization of SLPs’ reasoning processes for joint attention using interviews and video annotations.
    2. Development of an MLLM system achieving up to 85% accuracy in observation and 64% accuracy in judgment compared to expert labels.
    3. Creation of a dataset with expert-labeled joint attention behaviors and design guidelines for aligning MLLMs with SLP workflows.
  • Procedure and key techniques:
    1. Conduct interviews and annotation sessions with three SLPs to identify key behavioral dimensions (gaze, action, vocalization).
    2. Translate expert heuristics into structured prompts for MLLMs to extract behavioral segments from videos.
    3. Evaluate MLLM judgment accuracy using zero-shot and many-shot prompting strategies with GPT-4.1.

Results

  • Concrete findings:
    • Observation-level alignment: MLLMs achieved 81% accuracy in identifying gaze, action, and vocalization cues.
    • Judgment-level alignment: Many-shot prompting improved accuracy to 64%, macro-F1 to 0.44, and Cohen’s κ to 0.26.
  • Advantage over baselines: Structured prompts improved behavioral description accuracy from 62% to 81%, and judgment alignment outperformed zero-shot prompting.
  • Experiments / evaluation:
    • Dataset: 25 YouTube videos of parent–child interactions annotated by three SLPs.
    • Metrics: Accuracy, macro-precision, macro-recall, macro-F1, Cohen’s κ.
    • Models: Gemini 2.5 Pro for observation, GPT-4.1 for judgment.
  • Limitations and future work:
    • Small sample size of three SLPs limits generalizability.
    • Dataset skewed toward neurotypical interactions, with few examples of poor joint attention.
    • Stage 2 results based on manually corrected descriptions, representing an upper bound on performance.
    • Future work should expand datasets to include neurodiverse populations and larger expert cohorts.

Summary

This study explores the alignment of multimodal large language models (MLLMs) with speech-language pathologists (SLPs) in analyzing joint attention during parent–child interactions. The proposed two-stage system separates observation (behavioral cues like gaze, action, vocalization) from judgment (interaction quality assessment). Results show robust alignment at the observation level (81% accuracy) but modest alignment at the judgment level (64% accuracy with many-shot prompting). Findings highlight the feasibility of using MLLMs to support expert workflows and suggest design opportunities for human-centered AI systems that augment professional expertise in developmental assessments.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222624/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791267
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
2 authors
sell
Subtopics
Human-LLM Collaboration, Human Pose & Activity Recognition, Child-Computer Interaction Design
work
Professions
Speech-Language Pathologists & Audiologists, University Professors & Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers