This report delves into an explorative journey in the realm of reinforcement learning (RL), specifically applied to the classic game of tic-tac-toe in various dimensions. Our project was structured in three distinct phases, each higher in complexity and dimensionality. The first phase involved applying Value Iteration to a standard 2D 3
In this report, we candidly recount our journey, detailing the plethora of challenges, setbacks, and breakthroughs encountered along the way. It includes discussions of the strategies we considered, the ones we adopted, and those we discarded. The report is also an honest reflection on our failed attempts and the invaluable lessons learned from them. By sharing our exhaustive research process, experimentation, and methodical approach to problem-solving, we aim to provide insights into the practical applications of RL in progressively complex scenarios.
Keywords: Tic-tac-toe, Value Iteration, Online Learning, Q-Learning, DQN