PREDILECT: Preferences Delineated with Zero-Shot Language-based Reasoning in Reinforcement Learning
Authors
Preference-based reinforcement learning (RL) has emerged as a new field in robot learning, where humans play a pivotal role in shaping robot behavior by expressing preferences on different sequences of state-action pairs. However, formulating realistic policies for robots demands responses from humans to an extensive array of queries. In this work, we approach the sample-efficiency challenge by expanding the information collected per query to contain both preferences and optional text prompting. To accomplish this, we leverage the zero-shot capabilities of a large language model (LLM) to reason from the text provided by humans. To accommodate the additional query information, we reformulate the reward learning objectives to contain flexible highlights -- state-action pairs that contain relatively high information and are related to the features processed in a zero-shot fashion from a pretrained LLM. In both a simulated scenario and a user study, we reveal the effectiveness of our work by analyzing the feedback and its implications. Additionally, the collective feedback collected serves to train a robot on socially compliant trajectories in a simulated social navigation landscape. We provide video examples of the trained policies at https://sites.google.com/view/rl-predilect
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 83%
New Enactions of Expertise: Software Engineers’ Evaluation and Demonstration of Coding Expertise with AI Coding Assistants
CHI '26· Human-LLM Collaboration +2
- 83%
From Junior to Senior: Allocating Agency and Navigating Professional Growth in Agentic AI-Mediated Software Engineering
CHI '26· Human-LLM Collaboration +2
- 83%
Developer Interaction Patterns with Proactive AI: A Five-Day Field Study
IUI '26· AI-Assisted Decision-Making & Automation +2
- 71%
Vibe Coding Entanglements – Repositioning Boundaries of Intention, Authorship, and Responsibility in Programming with Generative AI
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
Gemini at Work: Knowledge Workers' Perceptions and Assessment of Productivity Gains
DIS '25· Generative AI (Text, Image, Music, Video) +3
- 71%
GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning
UIST '24· Generative AI (Text, Image, Music, Video) +2
- 71%
Exploring Passenger-Automated Vehicle Negotiation Utilizing Large Language Models for Natural Interaction
AutoUI '24· Automated Driving Interface & Takeover Design +2
- 67%
Discovering the Syntax and Strategies of Natural Language Programming with Generative Language Models
CHI '22· Generative AI (Text, Image, Music, Video) +1
- 67%
Comparing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication
CHI '23· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)