קובי גל

אקדמי בכיר

Optimizing Representations and Policies for Question Sequencing using Reinforcement Learning

Aqil Zainal Azhar, Avi Segal, Kobi Gal

This paper studies the use of Reinforcement Learning (RL) policies for optimizing the sequencing of online learning materials to students. Our approach provides an end to end pipeline for automatically deriving and evaluating robust representations of students’ interactions and policies for content sequencing in online educational settings. We conduct the training and evaluation offline based on a publicly available dataset of diverse student online activities used by tens of thousands of students. We study the influence of the state representations on the performance of the obtained policy and its robustness towards perturbations on the environment dynamics induced by stronger and weaker learners. We show that ‘bigger may not be better’, in that increasing the complexity of the state space does not necessarily lead to better performance, as measured by expected future reward. We describe two methods for offline evaluation of the policy based on importance sampling and Monte Carlo policy evaluation. This work is a first step towards optimizing representations when designing policies for sequencing educational content that can be used in the real world.

שפת פרסום אנגלית
סטטוס פרסום פורסם - 01.01.2022

ASJC Scopus subject areas

Artificial Intelligence
Computer Science Applications
Human-Computer Interaction
Information Systems
גישה למסמך
10.5281/zenodo.6853123
קבצים וקישורים אחרים
Link to publication in Scopus