
Yuval Shahar
A Deep Reinforcement Learning Framework based on Frequent Temporal Patterns as States for Optimal Therapy
Assessment in the Hypokalemia domain
Clinical Decision Support Systems (CDSSs) based on evidence-based clinical guidelines (GLs) enable real-time, consistent, and cost-effective medical decision-making. However, developing or modifying GLs remains a complex, expert-driven process. As a result, guidelines are often static, require local adaptation, and are not calibrated to institutional data or temporally complex patient trajectories, limiting their ability to generalize and evolve. We introduce TP-DRL, a framework that combines knowledge-based temporal abstraction, frequent temporal pattern mining, and conservative offline deep reinforcement learning to derive adaptive treatment policies directly from retrospective data. Applied to hypokalemia management using the MIMIC-IV ICU dataset, the pipeline transforms raw laboratory measurements, vital signs, and intervention sequences into temporally aware discrete dynamic states. A multi-horizon clinical reward function guides learning by jointly optimizing short-term biochemical correction, medium-term complication prevention, and long-term survival and discharge outcomes. The reward incorporates trajectory-efficiency penalties and treatment-balance constraints to discourage unnecessary interventions and looping behavior while ensuring safety. In patients clinically identified for potassium chloride treatment, use of TP-DRL policies would have potentially reduced mortality by 2%, increased discharge rates by 9.96%, and reduced the session duration by 28.1% while avoiding over-treatment. These results demonstrate that temporal-pattern-aware offline reinforcement learning can serve as a foundation for data-adaptive "living guidelines"that align with institutional practice, improve patient outcomes, and generalize to other complex clinical pathways.
| Publication language | English |
| Pages | 621-630 |
| Publication status | Published - 01.01.2026 |