Talking Spell: A Wearable System Enabling Real-Time Anthropomorphic Voice Interaction with Everyday Objects
Authors
Virtual assistants (VAs) have become ubiquitous in daily life, integrated into smartphones and smart devices, sparking interest in AI companions that enhance user experiences and foster emotional connections. However, existing companions are often embedded in specific objects—such as glasses, home assistants, or dolls—requiring users to form emotional bonds with unfamiliar items, which can lead to reduced engagement and feelings of detachment. To address this, we introduce Talking Spell, a wearable system that empowers users to imbue any everyday object with speech and anthropomorphic personas through a user-centric radiative network. Leveraging advanced computer vision (e.g., YOLOv11 for object detection), large vision-language models (e.g., QWEN-VL for persona generation), speech-to-text and text-to-speech technologies, Talking Spell guides users through three stages of emotional connection: acquaintance, familiarization, and bonding. We validated our system through a user study involving 12 participants, utilizing Talking Spell to explore four interaction intentions: entertainment, companionship, utility, and creativity. The results demonstrate its effectiveness in fostering meaningful interactions and emotional significance with everyday objects. Our findings indicate that Talking Spell creates engaging and personalized experiences, as demonstrated through various devices, ranging from accessories to essential wearables.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 67%
Emotional Dialogue Generation Using Image-Grounded Language Models
CHI '18· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
Understanding Affective Experiences with Conversational Agents
CHI '19· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
A New Uncanny Valley? The Effects of Speech Fidelity and Human Listener Gender on Social Perceptions of a Virtual-Human Speaker
CHI '22· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
Simulating Human Imprecision in Temporal Statements of Intelligent Virtual Agents
CHI '22· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
Super Kawaii Vocalics: Amplifying the “Cute” Factor in Computer Voice
CHI '25· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
Exploring Humor as a Repair Strategy During Communication Breakdowns with Voice Assistants
CUI '23· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
“Hello, This is a Voice Assistant Calling" When a Human Voice Calls Claiming to Be a Machine on an Ordinary Day
DIS '25· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
Typing Behavior is About More than Speed: Users' Strategies for Choosing Word Suggestions Despite Slower Typing Rates
MobileHCI '23· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
Agent-based Mediation on Smartphone Usage among Co-located Couples
MobileHCI '24· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)