This is a reimplementation of the AlphaZero algorithm as described by Silver et al. (2017).
We use PyTorch and Numpy as the primary backend.
The algorithm is applied to a custom game of repeated multi-dimensional TicTacToe, but could be easily adapted to play any other two-player perfect information game.
- PyTorch reimplementation of MCTS and policy training from scratch
- Custom game environment following the Gym API
- Parallel training workflow using Python's
multiprocessingmodule