Modeling Variation in Human Feedback with User Inputs: An Exploratory Methodology
Authors
To expedite the development process of interactive reinforcement learning (IntRL) algorithms, prior work often uses perfect oracles as simulated human teachers to furnish feedback signals. Those oracles typically derive from ground-truth knowledge or optimal policies, and provide dense and error-free feedback to a robot learner without delay. However, this machine-like feedback behavior fails to accurately represent the diverse patterns observed in human feedback, which may lead to unstable or unexpected algorithm performance in real-world human-robot interaction. To alleviate this limitation of oracles in oversimplifying user behavior, we propose a method for modeling variation in human feedback that can be applied to a standard oracle. We present a 5-dimensional model with 5 dimensions of feedback variation identified in prior work. This model enables the modification of feedback output from perfect oracles to introduce more human-like features. We demonstrate how each model attribute can impact on the learning performance of an IntRL algorithm through a simulation experiment. We also conduct a proof-of-concept study to illustrate how our model can be populated from people in two ways. The modeling results intuitively present the feedback variation among participants and help to explain the mismatch between oracles and human teachers. Overall, our method is a promising step towards refining simulated oracles by incorporating insights from real users.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 60%
Effects of Communication Directionality and AI Agent Differences in Human-AI Interaction
CHI '21· Human-LLM Collaboration +1
- 60%
Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance
CHI '21· Explainable AI (XAI) +1
- 60%
Model Sketching: Centering Concepts in Early-Stage Machine Learning Model Design
CHI '23· AI-Assisted Decision-Making & Automation +1
- 60%
Matching Mind and Method: Augmented Decision-Making with Digital Companions based on Regulatory Mode Theory
CHI '23· AI-Assisted Decision-Making & Automation
- 60%
One AI Does Not Fit All: A Cluster Analysis of the Laypeople’s Perception of AI Roles
CHI '23· Explainable AI (XAI) +1
- 60%
AI Knowledge: Improving AI Delegation through Human Enablement
CHI '23· Human-LLM Collaboration +1
- 60%
"Help Me Help the AI": Understanding How Explainability Can Support Human-AI Interaction
CHI '23· Explainable AI (XAI) +1
- 60%
Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making
CHI '25· Human-LLM Collaboration +1
- 60%
Editable XAI: Toward Bidirectional Human-AI Alignment with Co-Editable Explanations of Interpretable Attributes
CHI '26· Explainable AI (XAI) +1
- 60%
Emergent, not Immanent: A Baradian Reading of Explainable AI
CHI '26· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)