PREDILECT: Preferences Delineated with Zero-Shot Language-based Reasoning in Reinforcement Learning

Generative AI (Text, Image, Music, Video)Human-LLM CollaborationAI-Assisted Decision-Making & AutomationAutonomous Driving Engineers & Test DriversSoftware Engineers & DevelopersAI/ML Researchers & Engineers

Preference-based reinforcement learning (RL) has emerged as a new field in robot learning, where humans play a pivotal role in shaping robot behavior by expressing preferences on different sequences of state-action pairs. However, formulating realistic policies for robots demands responses from humans to an extensive array of queries. In this work, we approach the sample-efficiency challenge by expanding the information collected per query to contain both preferences and optional text prompting. To accomplish this, we leverage the zero-shot capabilities of a large language model (LLM) to reason from the text provided by humans. To accommodate the additional query information, we reformulate the reward learning objectives to contain flexible highlights -- state-action pairs that contain relatively high information and are related to the features processed in a zero-shot fashion from a pretrained LLM. In both a simulated scenario and a user study, we reveal the effectiveness of our work by analyzing the feedback and its implications. Additionally, the collective feedback collected serves to train a robot on socially compliant trajectories in a simulated social navigation landscape. We provide video examples of the trained policies at https://sites.google.com/view/rl-predilect

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/hri/140086/2024

AdRecommended

Learn AI Coding at CodeNow

At a Glance

Paper Snapshot

fact_check
dataset
Source
HRI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Human-LLM Collaboration, AI-Assisted Decision-Making & Automation
work
Professions
Autonomous Driving Engineers & Test Drivers, Software Engineers & Developers, AI/ML Researchers & Engineers
article
Content Status
Abstract only
hub
Related Papers
10 related papers