Lookee: Gaze Tracking-based Infant Vocabulary Comprehension Assessment and Analysis

Eye Tracking & Gaze InteractionSpecial Education TechnologyPhysicians, Nurses & CliniciansK-12 TeachersUniversity Professors & ResearchersEarly Childhood Educators

Research Background and Problem

  • Identified Problems or Challenges: Early identification of infants' vocabulary comprehension abilities is crucial for developmental assessment and intervention, but existing evaluation methods face significant challenges. For instance, parent-report methods are prone to subjective bias, while experimental methods such as the "Intermodal Preferential Looking Paradigm" (IPLP) require manual data collection and analysis, limiting their scalability in clinical and home settings. Additionally, there is a lack of objective and convenient tools for assessing language comprehension in infants.
  • Importance of the Problem: Infants begin to understand vocabulary between 9 and 12 months of age, a critical period for early detection of language delays and prediction of future language abilities. Accurate assessment of language skills at this stage can help children access appropriate interventions earlier and provide valuable data resources for language research.
  • Research Motivation and Related Work: Current widely used assessment tools, such as parent-report methods (e.g., M-B CDI) and experimental methods (e.g., IPLP), suffer from subjective bias and limitations in data reliability and scalability. Furthermore, while AI-driven methods have recently emerged, they have not yet been fully integrated into infant assessment tools.

Solution

  • Proposed Method or Solution: The authors designed an assessment tool called "Lookee," which leverages artificial intelligence and gaze-tracking technology to automate the evaluation of infants' vocabulary comprehension. Lookee combines the IPLP experimental method with a Random Forest model to analyze infants' gaze behavior, providing a more objective and efficient language assessment.
  • Innovative Aspects of the Solution:
    • Integration of AI technology with the IPLP experimental paradigm to offer an automated and scalable assessment solution.
    • Use of camera-based gaze tracking to reduce reliance on specialized equipment (e.g., eye trackers) and eliminate the need for complex calibration processes, simplifying the assessment procedure.
    • Support for personalized parameter configurations to meet the needs of diverse populations and research objectives.
  • Implementation Steps and Key Technologies:
    1. Experimental Design: A total of 40 vocabulary comprehension trials were set up, including target words and distractor words (presented as images).
    2. Data Collection and Analysis: Gaze point data were captured via a camera, and the Random Forest model was used to analyze gaze behavior, extracting features such as gaze proportion and reaction time.
    3. Result Reporting: A comprehensive evaluation report was generated, including age group comparison scores, vocabulary comprehension estimates, and detailed trial information.
    4. User Research: The tool's accuracy and practicality were validated through experimental studies and interviews, collecting feedback from various stakeholders (e.g., clinicians, parents, researchers).

Research Outcomes

  • Specific Outcomes:
    • In tests involving 19 parent-infant pairs, Lookee's analysis results showed over 80% accuracy compared to parent-reported vocabulary comprehension.
    • The tool scored 74.56 on the System Usability Scale (SUS), indicating good usability.
  • Advantages Over Existing Solutions:
    • Reduced the need for manual data recording and analysis, improving the efficiency and scalability of the assessment tool.
    • Provided an objective, automated evaluation method, effectively minimizing subjective bias in existing approaches.
    • Enabled remote assessments using simple equipment (e.g., standard webcams), making it convenient for use in home environments.
  • Experimental or Evaluation Results:
    • The Random Forest model demonstrated consistency in cross-validation, indicating its generalizability.
    • Higher accuracy (90.96%) was observed in older age groups (18–30 months), while younger groups (11–18 months) showed measurement errors likely due to difficulties in maintaining attention.
  • Limitations and Future Directions:
    • The dataset was relatively small, with only 31 prior experimental data points and 19 new experimental data points, limiting the model's generalizability.
    • The experimental setting was a children's museum, which may have introduced environmental interference; future studies should explore broader settings such as homes or clinical environments.
    • Current tests focused on a single cultural background in Korea; future expansions should include more languages, cultures, and clinical conditions (e.g., children with language delays or hearing impairments).
    • Developing more advanced AI models (e.g., LSTM and Transformer) and incorporating multimodal data inputs (e.g., facial expressions, gestures) could further enhance assessment accuracy and comprehensiveness.

The above analysis summarizes the main contributions and technological innovations of the study while highlighting its current limitations and potential future directions.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188210/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713386
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Special Education Technology
work
Professions
Physicians, Nurses & Clinicians, K-12 Teachers, University Professors & Researchers, Early Childhood Educators
article
Content Status
Full text indexed
hub
Related Papers
0 related papers