AI of Oz: Enhancing Wizard of Oz Studies in HCI with AI Assistance for Human Moderation

Human-LLM CollaborationUser Research Methods (Interviews, Surveys, Observation)Prototyping & User TestingHCI ResearchersAI/ML Researchers & Engineers

Paper Title

AI of Oz: Enhancing Wizard of Oz Studies in HCI with AI Assistance for Human Moderation

Publication Info

  • Topic area: AI-assisted frameworks for Wizard of Oz studies in Human-Computer Interaction (HCI).
  • Keywords: Wizard of Oz, Human-Computer Interaction, AI moderation, large language models, human-agent interaction, cognitive load, response latency, mental health, sensitivity detection, conversational agents.

Background and Problem

  • Problem / challenge: Traditional Wizard of Oz (WoZ) methods in HCI research impose high cognitive loads on human moderators, leading to response delays and difficulties in managing sensitive or complex interactions. Current AI solutions, such as large language models (LLMs), can generate biased, generic, or harmful content, making full automation undesirable.
  • Significance: Improving WoZ methodologies is critical for advancing HCI research, particularly in sensitive domains like mental health, where natural and safe interactions are essential.
  • Motivation and related work: Prior research has explored crowd-powered and semi-automated WoZ systems, but these approaches face challenges like high latency, inconsistency, and operational costs. Recent studies integrating LLMs into WoZ setups have shown promise but lack focus on supporting moderators’ workflows and decision-making.

Solution

  • Proposed approach: AI of Oz, an AI-assisted WoZ system that integrates LLMs to support human moderators by detecting sensitive inputs, generating response suggestions, and summarizing conversations in real time.
  • Novelty:
    1. Development of a modular AI-assisted WoZ system with detection, response, and summarization capabilities.
    2. Redefinition of the wizard’s role as a moderator rather than a sole content generator.
    3. Empirical evaluation of the system’s impact on workflow, cognitive load, and decision-making in a mental health context.
    4. Introduction of a human-on-the-loop (HOTL) collaboration model for scalable and ethical HAI research.
  • Procedure and key techniques:
    • Detection Module: Classifies user input into sensitivity levels (Green, Yellow, Red) to flag critical messages.
    • Response Module: Generates context-aware reply suggestions based on conversation history and predefined personas.
    • Summary Module: Provides real-time conversation summaries to help moderators maintain context.
    • System architecture includes a web-based interface with real-time monitoring, response review, and intervention capabilities.

Results

  • Concrete findings:
    • Response latency for sensitive messages reduced from 84.95 seconds (no AI) to 40.75 seconds (full AI support).
    • NASA-TLX scores showed significant reductions in mental workload, effort, and frustration when AI modules were active.
    • Trust ratings improved significantly with the Response Module, particularly in competence and benevolence dimensions.
    • Technology Acceptance Model scores indicated higher perceived usefulness, ease of use, and behavioral intention in the full AI condition.
  • Advantage over baselines:
    • The Response Module consistently reduced cognitive load and response latency compared to conditions without AI support.
    • Combined AI modules (detection and response) achieved the highest usability and trust ratings.
  • Experiments / evaluation:
    • A within-subjects study with 20 graduate researchers evaluated the system across four conditions (with/without detection and response modules).
    • Metrics included response latency, NASA-TLX workload scores, trust (HCTS), usability (SUS), and technology acceptance (TAM).
    • Role-playing scenarios focused on mental health conversations with scripted inputs and predefined stress levels.
  • Limitations and future work:
    • Role-playing setup may lack realism; future studies should involve real users and domain experts.
    • Limited comparison with other WoZ systems (e.g., crowd-powered approaches).
    • Focused on GPT-4o; future work could explore other advanced LLMs for improved performance.
    • Small sample size of graduate researchers; broader studies are needed for generalizability.

Summary

This paper introduces AI of Oz, an AI-assisted Wizard of Oz system that leverages large language models to support human moderators in HCI research. The system reduces cognitive load and response latency by detecting sensitive inputs, generating context-aware responses, and summarizing conversations in real time. A user study with 20 researchers demonstrated significant improvements in task efficiency, usability, and trust, particularly when the Response Module was active. While focused on mental health scenarios, the framework has potential applications in education, e-commerce, and other domains requiring human-AI collaboration. Future work will address limitations in realism, scalability, and model comparison to refine and expand the system’s applicability.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222411/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791324
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
9 authors
sell
Subtopics
Human-LLM Collaboration, User Research Methods (Interviews, Surveys, Observation), Prototyping & User Testing
work
Professions
HCI Researchers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers