Finished · MSc

Combining off and on-policy training in Deep Reinforcement

Authored by Alexandre João Gomes Borges

Supervised by Arlindo L. Oliveira

View the thesis →

MuZero is able to master both Atari games and board games by learning a model of the environment, that is then used with Monte Carlo Tree Search (MCTS) to decide what move to play in each position. During tree search, the algorithm simulates games by exploring several possible moves, and afterwards picks the action that corresponds to the most promising trajectory. Even though not all trajectories from these simulated games are useful, none of them are used for training. Using these trajectories would provide more data, more quickly, leading to faster convergence and sample efficiency. Recent work introduced an off-policy value target for AlphaZero that uses data from simulated games. Similarly, in this work, we propose a way to obtain off-policy targets by using data from simulated games in MuZero. We combine these off-policy targets with the on-policy targets already used in MuZero in several ways, and study the impact of these targets and their combinations in two environments with distinct characteristics.

← All dissertations