Quantifying Latencies: A Conversation Analysis Approach to Human-Agent Interactions in Virtual Reality
Authors
Paper Title
Quantifying Latencies: A Conversation Analysis Approach to Human-Agent Interactions in Virtual Reality
Publication Info
- Topic area: Human-agent interaction in virtual reality with a focus on conversational timing.
- Keywords: Latency, virtual reality, embodied conversational agents, human-agent interaction, turn-taking, speech-to-speech models, conversation analysis, overlaps, gaps, repair strategies.
Background and Problem
- Problem / challenge: Existing conversational agents, particularly those using large language models (LLMs), exhibit technological latencies that disrupt natural turn-taking, leading to user frustration and reduced adoption. Current studies lack empirical quantification of these latencies in human-agent interactions within VR environments.
- Significance: Understanding and addressing conversational latencies is critical for improving user experience, trust, and effectiveness in applications such as education, healthcare, and social VR.
- Motivation and related work: Prior research has explored latency effects in human-human and human-agent interactions but primarily through induced experimental conditions or subjective measures. There is a gap in analyzing naturalistic conversational timing in VR, particularly with speech-to-speech LLM agents.
Solution
- Proposed approach: A conversation analysis methodology to quantify latencies in human-agent interactions within VR, focusing on two metrics: Floor-Transfer Offset (FTO) and Post-Overlap Continuation (POC).
- Novelty:
- First empirical quantification of conversational latencies in VR-based human-agent interactions using speech-to-speech LLMs.
- Identification of two latency patterns: Start-Up Latencies (delayed agent responses) and Wind-Down Latencies (agent delays in stopping speech after interruptions).
- Development of design implications and heuristics for improving agent timing and responsiveness.
- Procedure and key techniques:
- Collection of dyadic interaction data in VRChat using Meta Quest 3 headsets.
- Extraction of speech features (silences, overlaps) using pyannote/speaker-diarization-3.1 and ELAN software.
- Quantification of FTO and POC metrics through statistical analysis and visualization.
- Annotation of transcripts using Jeffersonian transcription for detailed conversation analysis.
Results
- Concrete findings:
- Agent FTO durations (Mdn = 4050 ms) were significantly longer than user FTO durations (Mdn = 1232 ms), indicating Start-Up Latencies.
- Agent POC durations (Mdn = 931 ms) were significantly longer than user POC durations (Mdn = 287 ms), indicating Wind-Down Latencies.
- Overlap accounted for only 0.36% of conversation time, while silence accounted for 44.64%.
- Advantage over baselines:
- Empirical quantification of conversational timing metrics provides actionable insights beyond subjective measures used in prior studies.
- Identification of specific latency patterns enables targeted design improvements for ECAs.
- Experiments / evaluation:
- Dataset: 8 hours 56 minutes of dyadic recordings in VRChat across 20 participants.
- Metrics: FTO and POC durations, silence and overlap distributions.
- Statistical tests: Mann–Whitney U tests with effect sizes.
- Limitations and future work:
- Current methods are limited to dyadic interactions and do not account for technological latency effects such as network delays.
- User demographics (e.g., speaking anxiety, non-native English proficiency) may confound timing results.
- Future work should explore multi-speaker interactions, individual-level variances, and cross-modal comparisons.
Summary
This study introduces a novel framework for quantifying conversational latencies in human-agent interactions within VR environments. By analyzing Floor-Transfer Offset and Post-Overlap Continuation metrics, the authors identify two distinct latency patterns—Start-Up and Wind-Down Latencies—that disrupt natural turn-taking. Empirical findings reveal significant timing asymmetries between users and agents, highlighting areas for design improvement. The proposed metrics and methodologies offer valuable tools for evaluating and enhancing the realism of embodied conversational agents, with implications for applications in education, healthcare, and social VR.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
On the Intelligence and Knowledgeability of Virtual Agents
CHI '26· Social & Collaborative VR +2
- 71%
AI-exhibited Personality Traits Can Shape Human Self-concept through Conversations
CHI '26· Agent Personality & Anthropomorphism +2
- 71%
Cracking the Case Together: Role Perceptions in Human-AI Mystery Solving Dialogues
CHI '26· Human-LLM Collaboration +2
- 71%
Toward Natural and Companionable Virtual Agents via Cross-Temporal Emotional Modeling
CHI '26· Agent Personality & Anthropomorphism +2
- 67%
Mapping Machine Learning Advances from HCI Research to Reveal Starting Places for Design Innovation
CHI '18· Human-LLM Collaboration
- 63%
Empowered XR through Generative AI: Balancing Superpowers and Risks
CHI '26· Generative AI (Text, Image, Music, Video) +3
- 63%
FeelWave: Enabling Emotion-Aware Voice Interaction through Noise-Robust mmWave Emotion Sensing
CHI '26· Affective Human-Computer Dialogue +3
- 63%
“I Felt Bad After We Ignored Her”: Understanding How Interface-Driven Social Prominence Shapes Group Discussions with GenAI
CHI '26· Human-LLM Collaboration +3
- 63%
From Prompt to Presence: Co-Creating Personalised Emotional Sanctuaries in VR with Generative AI
IUI '26· Generative AI (Text, Image, Music, Video) +3
Based on Jaccard similarity of research subtopics & professions (≥60%)