רונן ברפמן

אקדמי בכיר

PTDRLHF

Parameter Tuning Using Deep Reinforcement Learning with Human Feedback

Elias Goldsztejn, Dan Rouven Suissa, Ronen Brafman

Many autonomous navigation algorithms require parameter re-tuning when facing new environments. This paper presents PTDRLHF, a parameter-tuning strategy that combines the Reinforcement Learning (RL)-based parameter tuning approach of Parameter Tuning using Deep Reinforcement Learning (PTDRL) [1] with human feedback to adaptively select from a predetermined set of parameters for a given navigation system in the context of social navigation. Our learning strategy is motivated by techniques for training language models using human feedback (HF) [2]. To the best of our knowledge, we are the first to implement an RLHF method for dynamic tuning in mobile navigation. In simulation, PTDRLHF preserves 20 % greater clearance from people and obstacles than PTDRL, with only a marginal decrease in average speed; in real-world trials, it outperforms the baseline on nearly all subjective evaluation measures.

שפת פרסום אנגלית
דפים 263-269
סטטוס פרסום פורסם - 01.01.2025

Keywords

Intelligent Robotics
Reinforcement Learning and Preference/Ranking

ASJC Scopus subject areas

Software
Artificial Intelligence
Computer Science Applications
קבצים וקישורים אחרים
Link to publication in Scopus