Gilad Katz

Senior Academic

Team of Rivals

Hierarchical Deep Reinforcement Learning and Behavior Cloning for Multiplayer Poker

Avishag Shapira, Ido Rom, Asaf Shabtai,Gilad Katz

Multiplayer no-limit Texas Hold’em is considered a challenging benchmark for AI algorithms, due to the need for decision making under partial information, strategic deception, and non-stationary opponents. Classical equilibrium-based techniques do not extend cleanly to the multiplayer setting, and prevailing multiplayer solutions, such as LLMs, tend to be computationally intensive. This study introduces Havoc, a hierarchical deep RL approach that combines behavior cloning of individual human experts with a value-based master policy that selects, at each decision point, which specialist to deploy. By preserving distinct human play styles in the specialist policies and learning when to deploy them, Havoc adapts its strategy rapidly as table conditions shift. Despite limited training data, Havoc attains strong multiplayer performance, outperforming current state-of-the-art methods while also requiring substantially less computational resources.

Publication language English
Pages 2473-2481
Publication status Published - 24.05.2026

Keywords

behavioral cloning
deep reinforcement learning
poker

ASJC Scopus subject areas

Artificial Intelligence
Access to Document
10.65109/LRNX6318
Other files and links
Link to publication in Scopus