AI of Oz: Enhancing Wizard of Oz Studies in HCI with AI Assistance for Human Moderation
Authors
Paper Title
AI of Oz: Enhancing Wizard of Oz Studies in HCI with AI Assistance for Human Moderation
Publication Info
- Topic area: AI-assisted frameworks for Wizard of Oz studies in Human-Computer Interaction (HCI).
- Keywords: Wizard of Oz, Human-Computer Interaction, AI moderation, large language models, human-agent interaction, cognitive load, response latency, mental health, sensitivity detection, conversational agents.
Background and Problem
- Problem / challenge: Traditional Wizard of Oz (WoZ) methods in HCI research impose high cognitive loads on human moderators, leading to response delays and difficulties in managing sensitive or complex interactions. Current AI solutions, such as large language models (LLMs), can generate biased, generic, or harmful content, making full automation undesirable.
- Significance: Improving WoZ methodologies is critical for advancing HCI research, particularly in sensitive domains like mental health, where natural and safe interactions are essential.
- Motivation and related work: Prior research has explored crowd-powered and semi-automated WoZ systems, but these approaches face challenges like high latency, inconsistency, and operational costs. Recent studies integrating LLMs into WoZ setups have shown promise but lack focus on supporting moderators’ workflows and decision-making.
Solution
- Proposed approach: AI of Oz, an AI-assisted WoZ system that integrates LLMs to support human moderators by detecting sensitive inputs, generating response suggestions, and summarizing conversations in real time.
- Novelty:
- Development of a modular AI-assisted WoZ system with detection, response, and summarization capabilities.
- Redefinition of the wizard’s role as a moderator rather than a sole content generator.
- Empirical evaluation of the system’s impact on workflow, cognitive load, and decision-making in a mental health context.
- Introduction of a human-on-the-loop (HOTL) collaboration model for scalable and ethical HAI research.
- Procedure and key techniques:
- Detection Module: Classifies user input into sensitivity levels (Green, Yellow, Red) to flag critical messages.
- Response Module: Generates context-aware reply suggestions based on conversation history and predefined personas.
- Summary Module: Provides real-time conversation summaries to help moderators maintain context.
- System architecture includes a web-based interface with real-time monitoring, response review, and intervention capabilities.
Results
- Concrete findings:
- Response latency for sensitive messages reduced from 84.95 seconds (no AI) to 40.75 seconds (full AI support).
- NASA-TLX scores showed significant reductions in mental workload, effort, and frustration when AI modules were active.
- Trust ratings improved significantly with the Response Module, particularly in competence and benevolence dimensions.
- Technology Acceptance Model scores indicated higher perceived usefulness, ease of use, and behavioral intention in the full AI condition.
- Advantage over baselines:
- The Response Module consistently reduced cognitive load and response latency compared to conditions without AI support.
- Combined AI modules (detection and response) achieved the highest usability and trust ratings.
- Experiments / evaluation:
- A within-subjects study with 20 graduate researchers evaluated the system across four conditions (with/without detection and response modules).
- Metrics included response latency, NASA-TLX workload scores, trust (HCTS), usability (SUS), and technology acceptance (TAM).
- Role-playing scenarios focused on mental health conversations with scripted inputs and predefined stress levels.
- Limitations and future work:
- Role-playing setup may lack realism; future studies should involve real users and domain experts.
- Limited comparison with other WoZ systems (e.g., crowd-powered approaches).
- Focused on GPT-4o; future work could explore other advanced LLMs for improved performance.
- Small sample size of graduate researchers; broader studies are needed for generalizability.
Summary
This paper introduces AI of Oz, an AI-assisted Wizard of Oz system that leverages large language models to support human moderators in HCI research. The system reduces cognitive load and response latency by detecting sensitive inputs, generating context-aware responses, and summarizing conversations in real time. A user study with 20 researchers demonstrated significant improvements in task efficiency, usability, and trust, particularly when the Response Module was active. While focused on mental health scenarios, the framework has potential applications in education, e-commerce, and other domains requiring human-AI collaboration. Future work will address limitations in realism, scalability, and model comparison to refine and expand the system’s applicability.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Mapping the Wizards' Path: A Systematic Review of Wizard-of-Oz in HCI
CHI '26· Participatory Design +3
- 83%
Evaluating Generative AI in the Lab: Methodological Challenges and Guidelines
IUI '26· Generative AI (Text, Image, Music, Video) +3
- 83%
From Narrative to Numbers: Evaluating Survey Questionnaires with Large Language Models
IUI '26· Human-LLM Collaboration +2
- 83%
Rationalizer: Leveraging LLM to Support User Providing the Rationales Behind the Rating of Likert Scale Questionnaires
IUI '26· Human-LLM Collaboration +2
- 71%
DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models
CHI '26· Human-LLM Collaboration +3
- 71%
Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
CHI '26· Human-LLM Collaboration +3
- 71%
Evalet: Evaluating Large Language Models through Functional Fragmentation
CHI '26· Human-LLM Collaboration +3
- 71%
Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
CHI '26· Human-LLM Collaboration +3
- 71%
Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
CHI '26· Human-LLM Collaboration +3
- 71%
Live in the Loop: Rapid Run-time Feedback for Prompts
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)