גלעד כץ

אקדמי בכיר

Team of Rivals

Hierarchical Deep Reinforcement Learning and Behavior Cloning for Multiplayer Poker

Avishag Shapira, Ido Rom, Asaf Shabtai,Gilad Katz

Multiplayer no-limit Texas Hold’em is considered a challenging benchmark for AI algorithms, due to the need for decision making under partial information, strategic deception, and non-stationary opponents. Classical equilibrium-based techniques do not extend cleanly to the multiplayer setting, and prevailing multiplayer solutions, such as LLMs, tend to be computationally intensive. This study introduces Havoc, a hierarchical deep RL approach that combines behavior cloning of individual human experts with a value-based master policy that selects, at each decision point, which specialist to deploy. By preserving distinct human play styles in the specialist policies and learning when to deploy them, Havoc adapts its strategy rapidly as table conditions shift. Despite limited training data, Havoc attains strong multiplayer performance, outperforming current state-of-the-art methods while also requiring substantially less computational resources.

שפת פרסום אנגלית
דפים 2473-2481
סטטוס פרסום פורסם - 24.05.2026

Keywords

behavioral cloning
deep reinforcement learning
poker

ASJC Scopus subject areas

Artificial Intelligence
גישה למסמך
10.65109/LRNX6318
קבצים וקישורים אחרים
Link to publication in Scopus