Bayesian Reinforcement Learning for the n-Pile Nim Game
Abstract
Machine learning and reinforcement learning methods have achieved strong empirical performance across many applications, including game playing. However, this also raises concerns about the reliability, interpretability, and safety of model decisions, thereby motivating the study of learning algorithms in settings where the underlying problem structure is mathematically understood. Combinatorial games provide a useful setting for studying these questions because many admit mathematically characterized optimal strategies. Therefore, this paper uses the n-pile Nim game as a mathematically structured environment for studying Bayesian reinforcement learning through random-walk Metropolis Markov Chain Monte Carlo (MCMC) methods. The paper investigates how probabilistic posterior sampling over policy parameters influences the learned strategy and whether the resulting policy can recover the known combinatorial structure of optimal play. By combining theoretical analysis and empirical experiments, the paper shows that the proposed Bayesian reinforcement learning framework is able to recover the combinatorial structure underlying optimal play in the Nim game and achieve strong gameplay performance across multiple game settings.
Results