> ML_LITERATURE // SILVER-2017-MASTERING-CHESS-SHOGI-GO-SELF-PLAY-ALPHAZERO_v1.0
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm (AlphaZero)
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, Demis Hassabis · Science (2017)
seminal-architecture2017foundationalthirdPartyReproduced
Principal Contribution
Mastered chess, shogi, and Go from scratch within 24 hours starting from random play with zero human domain knowledge or opening books, using pure MCTS self-play.
Operational Relevance
Serves as qualified theoretical and systems foundation for task-game-playing.
Assumptions
- Markovian state dynamics and stationary reward functions hold in target evaluation environments
Limitations
- Sample efficiency, exploration stability, and real-world sim-to-real transfer gaps require specialized tuning
