Peeking Ahead of the Field Study: Exploring VLM Personas as Support Tools for Embodied Studies in HCI
Honorable MentionAuthors
Paper Title
Peeking Ahead of the Field Study: Exploring VLM Personas as Support Tools for Embodied Studies in HCI
Publication Info
- Topic area: Use of Vision-Language Models (VLMs) as proxies for human participants in field studies, particularly in Human-Computer Interaction (HCI).
- Keywords: Vision-Language Models, field studies, autonomous vehicles, pedestrian interaction, HCI research, eHMI, simulation, embodied interaction, AI personas, behavioral mimicry.
Background and Problem
- Problem / challenge: Field studies are essential for understanding human behavior in real-world contexts but are costly, time-intensive, and prone to errors. Current alternatives, such as mixed reality and AI agents, fail to fully replicate natural human responses or provide reliable behavioral variability.
- Significance: Reducing the cost, time, and risks associated with field studies while maintaining their ecological validity could significantly benefit HCI research and applications, such as autonomous vehicle (AV)-pedestrian interaction.
- Motivation and related work: Prior research has explored AI personas and mixed reality for simulating human behavior, but these approaches lack validation against real-world human responses. The study addresses this gap by systematically comparing VLM personas to human participants in a high-fidelity AV-pedestrian interaction task.
Solution
- Proposed approach: Use Vision-Language Model (VLM) personas to simulate human-like responses in field studies, focusing on AV-pedestrian interactions.
- Novelty:
- Introduced VLM personas to simulate complex human behaviors and spatial perception in embodied tasks.
- Conducted parallel experiments comparing real human participants and VLM personas in a controlled AV-pedestrian interaction scenario.
- Provided guidelines for using VLM personas in HCI research, emphasizing their strengths and limitations.
- Explored broader applications of VLM personas in embodied interaction beyond AV-pedestrian tasks.
- Procedure and key techniques:
- Conducted two parallel studies: a field study with 20 human participants and a video-based study with 20 VLM personas.
- Developed a questionnaire-based protocol to construct VLM personas mimicking human participants.
- Designed a video simulator with discretized spatial-temporal grids to replicate real-world scenarios for VLM personas.
- Compared behavioral and subjective data between humans and VLMs using metrics like crossing time, subjective trust, and trajectory features.
- Interviewed five HCI researchers to assess the applicability of VLM personas in research workflows.
Results
- Concrete findings:
- Crossing times were similar between humans (5.07s, SD=1.67) and VLM personas (5.25s, SD=0.72), with no significant difference (p=0.8465).
- VLM personas exhibited reduced behavioral variability compared to humans.
- Subjective trust ratings were comparable (humans: 3.03, VLMs: 2.97), but VLMs showed inflated confidence ratings (VLM: 4.53 vs. human: 3.50).
- VLMs captured consensus-level reasoning but missed minority perspectives and nuanced interpretations.
- Advantage over baselines:
- VLM personas effectively mimicked average human behavior and subjective perceptions, providing a low-cost alternative for exploratory and pilot studies.
- They reduced the need for extensive field study preparation while maintaining alignment with human data.
- Experiments / evaluation:
- Conducted a three-step process: field study, video simulation with VLMs, and expert interviews.
- Metrics included crossing time, subjective trust/confidence, trajectory features, and thematic analysis of open-ended responses.
- Experimental conditions included three eHMI types (light strip, animated eyes, no eHMI) and two AV behaviors (stop, non-stop).
- Limitations and future work:
- VLM personas lacked behavioral variability and failed to capture minority perspectives.
- Current VLMs lack binocular vision and depth perception, limiting their spatial reasoning capabilities.
- Future work should explore multimodal input (e.g., audio, haptics) and validate VLMs in broader embodied interaction domains.
Summary
This study demonstrates the potential of Vision-Language Model (VLM) personas to simulate human behavior in field studies, particularly in AV-pedestrian interaction scenarios. By comparing VLM personas to real human participants, the study found that VLMs can mimic average human responses but lack variability and depth. Three guidelines were proposed for using VLM personas in HCI research, emphasizing their suitability for exploratory and pilot studies but cautioning against over-reliance for nuanced insights. Interviews with HCI researchers highlighted potential applications, such as formative studies, large-scale simulations, and research involving underrepresented groups. Future work should address technical limitations and explore broader applications in embodied interaction.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)