How Do We Evaluate Experiences in Immersive Environments?
Authors
Paper Title
How Do We Evaluate Experiences in Immersive Environments?
Publication Info
- Topic area: Evaluation methodologies for immersive environments in HCI and XR research.
- Keywords: immersive experience, evaluation methods, virtual reality, augmented reality, questionnaires, task performance, physiological measures, user-centered design, computational modeling, open science.
Background and Problem
- Problem / challenge: Immersive experience evaluation is fragmented, with inconsistent methods, overlapping constructs, and limited comparability across studies. Core constructs like presence and embodiment are defined and measured in conflicting ways.
- Significance: Reliable evaluation of immersive experiences is essential for validating and improving interactive systems in domains like gaming, health, education, and collaboration.
- Motivation and related work: Prior work has focused on specific constructs, methods, or technologies, but lacks a comprehensive synthesis of evaluation practices across domains and decades. Foundational textbooks and reviews have highlighted the richness of the field but failed to unify its diverse perspectives.
Solution
- Proposed approach: A bottom-up scoping review of 375 papers from seven premier venues (ACM CHI, UIST, VRST, SUI, IEEE VR, ISMAR, TVCG) to map evaluation practices and provide a forward-looking agenda for immersive experience research.
- Novelty:
- Empirical synthesis of immersive experience evaluation practices across technologies and domains.
- Identification of recurring patterns, inconsistencies, and gaps in evaluation methods.
- Proposal of domain-sensitive, user-centered, and computationally integrated approaches for future evaluation.
- Advocacy for open and sustainable evaluation infrastructures.
- Procedure and key techniques:
- Literature search using keyword queries in ACM Digital Library and IEEE Xplore.
- Screening and coding of papers based on inclusion/exclusion criteria.
- Classification of papers by device modality, contribution type, application domain, and evaluation methods.
- Analysis of trends in method adoption, questionnaire usage, and methodological combinations.
Results
- Concrete findings:
- Questionnaires dominate evaluation methods (N = 321, 85.6%), followed by task performance (N = 235, 62.7%), interviews (N = 169, 45.1%), system performance metrics (N = 52, 13.9%), and physiological measures (N = 40, 10.7%).
- Multi-method evaluations are increasingly common, with 42.1% of papers using two methods and 32.3% using three.
- Questionnaire usage has risen over time, with recent studies employing up to eight instruments.
- Advantage over baselines:
- Comprehensive mapping of evaluation practices across domains, devices, and decades.
- Identification of methodological gaps and opportunities for integration.
- Emphasis on smarter, context-sensitive combinations of methods rather than redundant accumulation.
- Experiments / evaluation:
- Analysis of 375 papers published between 1995 and 2024, categorized by venue, device type, contribution type, and application domain.
- Evaluation methods classified into five categories: questionnaires, task performance, system performance, interviews, and physiological measures.
- Limitations and future work:
- Dataset excludes some venues and interdisciplinary perspectives.
- Keyword-based retrieval may miss relevant papers with divergent terminology.
- Future work should integrate cross-disciplinary insights, computational modeling, and open science infrastructures.
Summary
This paper provides a comprehensive review of immersive experience evaluation practices, synthesizing data from 375 papers across seven major venues. It highlights the fragmentation of methods, constructs, and metrics, emphasizing the need for domain-sensitive, user-centered, and computationally integrated approaches. Key findings include the dominance of questionnaires, the rise of multi-method evaluations, and the methodological skew toward feasibility over conceptual appropriateness. The authors advocate for smarter combinations of evaluation methods and the development of open, sustainable infrastructures to support reproducibility and cumulative progress in immersive experience research.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 75%
Verifying Finger-Fitts Models for Normalizing Subjective Speed-Accuracy Biases
MobileHCI '24· Prototyping & User Testing +1
- 67%
AutoStructGUI: Bridging Design and Implementation of GUI through Structured Layout Generation
IUI '26· Computational Methods in HCI +2
- 60%
Touchstone2: An Interactive Environment for Exploring Trade-offs in HCI Experiment Design
CHI '19· User Research Methods (Interviews, Surveys, Observation) +2
- 60%
How Users Interpret Bugs in Trigger-Action Programming
CHI '19· Prototyping & User Testing +1
- 60%
Modeling Fully and Partially Constrained Lasso Movements in a Grid of Icons
CHI '19· Prototyping & User Testing +1
- 60%
ORCSolver: An Efficient Solver for Adaptive GUI Layout with OR-Constraints
CHI '20· Prototyping & User Testing +1
- 60%
Investigating the Homogenization of Web Design: A Mixed-Methods Approach
CHI '21· Prototyping & User Testing +1
- 60%
How to Evaluate Object Selection and Manipulation in VR? Guidelines from 20 Years of Studies
CHI '21· Immersion & Presence Research +1
- 60%
Exploring Technical Reasoning in Digital Tool Use
CHI '22· Prototyping & User Testing +1
- 60%
i-LaTeX: Manipulating Transitional Representations between LaTeX Code and Generated Documents
CHI '22· Prototyping & User Testing +1
Based on Jaccard similarity of research subtopics & professions (≥60%)