Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?
Authors
Large Language Models have shown exceptional generative abilities in various natural language and generation tasks. However, possible anthropomorphization and leniency towards failure cases have propelled discussions on emergent abilities of Large Language Models especially on Theory of Mind (ToM) abilities in Large Language Models. While several false-belief tests exists to verify the ability to infer and maintain mental models of another entity, we study a special application of ToM abilities that has higher stakes and possibly irreversible consequences : Human Robot Interaction. In this work, we explore the task of Perceived Behavior Recognition, where a robot employs a Large Language Model (LLM) to assess the robot's generated behavior in a manner similar to human observer. We focus on four behavior types, namely - explicable, legible, predictable, and obfuscatory behavior which have been extensively used to synthesize interpretable robot behaviors. The LLMs goal is, therefore to be a human proxy to the agent, and to answer how a certain agent behavior would be perceived by the human in the loop, for example "Given a robot's behavior X, would the human observer find it explicable?". We conduct a human subject study to verify that the users are able to correctly answer such a question in the curated situations (robot setting and plan) across five domains. A first analysis of the belief test yields extremely positive results inflating ones expectations of LLMs possessing ToM abilities. We then propose and perform a suite of perturbation tests which breaks this illusion, i.e. Inconsistent Belief, Uninformative Context and Conviction Test. We conclude that, the high score of LLMs on vanilla prompts showcases its potential use in HRI settings, however to possess ToM demands invariance to trivial or irrelevant perturbations in the context which LLMs lack. We report our results on GPT-4 and GPT-3.5-turbo.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 63%
Emotion-aware Design in Automobiles: Embracing Technology Advancements to Enhance Human-vehicle Interaction
CHI '25· Brain-Computer Interface (BCI) & Neurofeedback +1
- 63%
Effects of LLM-based Search on Decision Making: Speed, Accuracy, and Overreliance
CHI '25· Human-LLM Collaboration +2
- 63%
Understanding Socio-technical Factors Configuring AI Non-Use in UX Work Practices
CHI '25· Human-LLM Collaboration +2
- 63%
Exploring The Impact of Proactive Generative AI Agent Roles In Time-Sensitive Collaborative Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 63%
Does My Chatbot Have an Agenda? Understanding Human and AI Agency in Human-Human-like Chatbot Interaction
CHI '26· Agent Personality & Anthropomorphism +2
- 63%
Towards AI as Colleagues: Multi-Agent System Improves Structured Ideation Processes
CHI '26· Human-LLM Collaboration +2
- 63%
Situated, Dynamic, and Subjective: Envisioning the Design of Theory-of-Mind-Enabled Everyday AI with Industry Practitioners
CHI '26· Brain-Computer Interface (BCI) & Neurofeedback +2
- 63%
DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented Interfaces
CHI '26· Human-LLM Collaboration +2
- 63%
Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks
CHI '26· Agent Personality & Anthropomorphism +2
- 63%
Machine Eye: Designing Relational Engagement with Embodied Large Language Models
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)