רונן ברפמן

אקדמי בכיר

Reinforcement Learning n RDPs by Combining Deep RL with Automata Learning

Tal Shahar, Ronen I. Brafman

Regular Decision Processes (RDPs) are a recently introduced model for decision-making in non-Markovian domains in which states are not postulated a-priori, and the next observation depends in a regular manner on past history. As such, they provide a more succinct and understandable model of the dynamics and reward function. Existing algorithms for learning RDPs attempt to learn an automaton that reflects the regularity of the underlying domain. However, their scalability is limited due to the practical difficulty of learning automata. In this paper we propose to leverage the power of Deep reinforcement learning in partially observable domain to learn RDPs: First, we learn an RNN-based policy. Then, we generate an automaton that reflects the policy's structure and use our old data to transform it into an MDP, which we solve. This results in a finite, explainable policy structure, and, as our empirical evaluation on old and new RDP benchmarks shows, much better sample complexity.

שפת פרסום אנגלית
דפים 2097-2104
סטטוס פרסום פורסם - 28.09.2023

ASJC Scopus subject areas

Artificial Intelligence
גישה למסמך
10.3233/FAIA230504
קבצים וקישורים אחרים
Link to publication in Scopus