How Do We Research Human-Robot Interaction in the Age of Large Language Models? A Systematic Review
Authors
Paper Title
How Do We Research Human-Robot Interaction in the Age of Large Language Models? A Systematic Review
Publication Info
- Topic area: Systematic review of the integration of large language models (LLMs) into human-robot interaction (HRI).
- Keywords: Large language models, human-robot interaction, multimodal perception, generative interaction, alignment, taxonomy, evaluation metrics, robotics.
Background and Problem
- Problem / challenge: Existing research on HRI has focused on technical aspects of LLMs, such as robustness and architecture, but lacks a comprehensive synthesis of their human-centered implications, including user modeling, emotional understanding, and levels of autonomy.
- Significance: Understanding how LLMs transform HRI is critical for advancing embodied intelligence and enabling seamless human-robot collaboration in real-world settings.
- Motivation and related work: Prior reviews have analyzed LLMs in robotics, focusing on technical advancements and specific applications. However, they fail to address the broader human-centered challenges and opportunities that arise from integrating LLMs into HRI. This paper aims to fill this gap by systematically reviewing and categorizing the field.
Solution
- Proposed approach: A systematic review of 86 studies on LLM-driven HRI, structured around a Sense–Interaction–Alignment framework.
- Novelty:
- Introduction of the Sense–Interaction–Alignment framework to conceptualize LLM-driven HRI.
- Development of a taxonomy categorizing research across nine dimensions, including perception, interaction, and evaluation.
- Identification of 11 key challenges for future research, such as multimodal grounding, trust calibration, and long-term alignment.
- Creation of an open-access database of the reviewed studies for transparency and reproducibility.
- Procedure and key techniques:
- Literature search following PRISMA guidelines across ACM DL, IEEE Xplore, and other databases.
- Inclusion criteria focused on studies integrating LLMs into embodied HRI systems.
- Analysis structured around research questions addressing foundational capabilities, system design, evaluation strategies, and future challenges.
Results
- Concrete findings:
- LLMs enhance HRI by enabling contextual sensing, generative interaction, and adaptive alignment.
- Research is highly heterogeneous, with diverse experimental setups, modalities, and evaluation metrics.
- Applications span eight domains, including healthcare, education, and industrial manufacturing.
- Advantage over baselines:
- LLMs improve task efficiency (e.g., reducing interaction time by up to 50%), enhance social interaction through adaptive dialogue, and enable multimodal reasoning.
- Experiments / evaluation:
- Methodologies include laboratory experiments (58 studies), field deployments (17 studies), and user-centered evaluations (e.g., interviews, questionnaires).
- Evaluation metrics cover task efficiency, accuracy, perceived intelligence, and user satisfaction.
- Limitations and future work:
- Challenges include multimodal alignment, emotional intelligence, trust calibration, and privacy risks.
- Future work should focus on longitudinal studies, proactive repair mechanisms, and ethical considerations in personalization.
Summary
This systematic review synthesizes how large language models (LLMs) are transforming human-robot interaction (HRI) by enabling advanced contextual sensing, generative interaction, and adaptive alignment. The proposed Sense–Interaction–Alignment framework organizes these capabilities and highlights key design considerations and challenges, such as multimodal grounding and trust calibration. The findings are supported by a comprehensive analysis of 86 studies, covering diverse applications and evaluation strategies. This work provides a roadmap for future research to address the technical and ethical complexities of LLM-driven HRI.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)