Papers
Topics
Authors
Recent
Search
2000 character limit reached

Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Published 5 Dec 2017 in cs.AI and cs.LG | (1712.01815v1)

Abstract: The game of chess is the most widely-studied domain in the history of artificial intelligence. The strongest programs are based on a combination of sophisticated search techniques, domain-specific adaptations, and handcrafted evaluation functions that have been refined by human experts over several decades. In contrast, the AlphaGo Zero program recently achieved superhuman performance in the game of Go, by tabula rasa reinforcement learning from games of self-play. In this paper, we generalise this approach into a single AlphaZero algorithm that can achieve, tabula rasa, superhuman performance in many challenging domains. Starting from random play, and given no domain knowledge except the game rules, AlphaZero achieved within 24 hours a superhuman level of play in the games of chess and shogi (Japanese chess) as well as Go, and convincingly defeated a world-champion program in each case.

Citations (1,625)

Summary

  • The paper demonstrates that a generalized reinforcement learning algorithm, AlphaZero, can quickly surpass elite chess and shogi programs through self-play refinement.
  • It uniquely combines deep neural networks with Monte-Carlo Tree Search to optimize move selection while reducing the number of evaluated positions compared to conventional methods.
  • Experimental results show that AlphaZero defeats world-champion engines like Stockfish and Elmo within hours, underscoring its computational efficiency and broad applicability.

Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

This paper presents AlphaZero, a generalized reinforcement learning algorithm designed to achieve superhuman performance across multiple complex domains without any domain-specific knowledge, except for the game rules. The algorithm builds upon the success of AlphaGo Zero, which demonstrated remarkable performance in the game of Go. AlphaZero extends this capability to other strategic games, namely chess and shogi.

Algorithm Overview

AlphaZero synthesizes deep neural networks with Monte-Carlo Tree Search (MCTS) to efficiently explore vast state spaces characteristic of chess, shogi, and Go. The neural network, parameterized by θ\theta, outputs the move probabilities p\mathbf{p} and state value vv for any given position ss. Through self-play, AlphaZero iteratively improves its policy (pθ)(\mathbf{p}_\theta) and value estimates (vθ)(v_\theta), facilitating its MCTS to conduct more targeted and efficient searches.

Training and Evaluation

AlphaZero’s training paradigm involves a robust regime of self-play and reinforcement learning, fulfilling several key milestones:

  • Chess: Outperformed the TCEC 2016 world-champion program Stockfish after four hours of training.
  • Shogi: Surpassed the 2017 CSA world-champion program Elmo in less than two hours.
  • Go: Demonstrated superiority over AlphaGo Lee after eight hours, achieving this with a fraction of the computational resources and time used in prior models.

Experimental Results

Detailed results from evaluation matches highlight the efficacy of AlphaZero:

  • Chess: In matches comprising 100 games at standard tournament time controls (one minute per move), AlphaZero defeated Stockfish convincingly, losing zero games and drawing or winning the remaining matches.
  • Shogi: Exhibited strong performance against Elmo, suffering only minimal losses.
  • Go: Achieved consistent victories against prior versions of AlphaGo Zero, solidifying its generalization capabilities.

Figure~\ref{fig:training} and Table~\ref{tab:results} provide quantitative insights into the performance trajectory of AlphaZero during training and the outcomes of the evaluation matches.

Computational Efficiency

AlphaZero’s efficiency is notable, as its MCTS evaluates significantly fewer positions per second compared to traditional alpha-beta search engines yet achieves superior performance. The search extends over critical lines of play through selective deep dives facilitated by the neural network’s policy and value estimates. This contrasts starkly with engines like Stockfish and Elmo, which rely on exhaustive search spaces and human-crafted heuristics.

Implications and Future Directions

Practically, AlphaZero's ability to master multiple strategy games from scratch showcases the potential for generalized algorithms in varied domains. Theoretically, the results challenge traditional beliefs regarding the supremacy of alpha-beta search in strategic games, positing Monte-Carlo methods augmented with neural networks as a viable and often superior alternative.

Future developments in AI might expand upon this framework to tackle real-time decision-making tasks and complex simulations beyond board games. Further investigations could integrate domain-specific tweaks or multi-domain learning capabilities to further enhance AlphaZero’s adaptability and performance.

Conclusion

AlphaZero is a significant advancement in the application of reinforcement learning to complex strategy games. By eschewing domain-specific knowledge and employing a unified approach to learning, it transcends the limitations of traditional game-specific algorithms, pointing toward a new horizon in the development of general AI systems. The convergence of deep learning and MCTS augurs well for applications requiring strategic planning and real-time decision-making, holding promise for diverse and impactful AI-driven innovations.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 19 tweets with 2684 likes about this paper.

YouTube

Alpha Zero's "Immortal Zugzwang Game" against Stockfish 1.7M views
Google's self-learning AI AlphaZero masters chess in 4 hours 1.6M views
Google Deep Mind AI Alpha Zero Refutes 1.e4 1.1M views
Google Deep Mind AI Alpha Zero Devours Stockfish 1.1M views
Google Deep Mind Alpha Zero Sacs a Piece Without "Thinking" Twice 885K views
Deep Mind AI Alpha Zero Sacrifices a Pawn and Cripples Stockfish for the Entire Game 542K views
Deep Mind AI Alpha Zero Dismantles Stockfish's French Defense 374K views
Deep Mind AI Alpha Zero Refuses a Draw from Stockfish 366K views
Deep Mind AI Alpha Zero's Positional Masterpiece With the Black Pieces 347K views
AlphaZero stuns with another brilliant move Engines took hours to understand 264K views
AlphaZero from Scratch – Machine Learning Tutorial 173K views
AlphaZero - Stockfish: ФАНТАСТИЧЕСКАЯ ПАРТИЯ во французской защите! 124K views
AlphaZéro l'intelligence artificielle qui met les échecs en PLS 100K views
AlphaZero: DeepMind's New Chess AI | Two Minute Papers #216 88K views
Ajedrez e informática: Desmitificando AlphaZero vs. Stockfish - Entrevista al Dr. Juanjo del Coz 85K views
Живые татуировки из бактерий и секрет жирной диеты в главных новостях на QWERTY. 75K views
Outrageous Artificial Intelligence: (Game 1) DeepMind’s AlphaZero crushes Stockfish Chess WC 72K views
РЕВОЛЮЦИЯ в шахматах! Новый алгоритм AlphaZero победил Stockfish! 66K views
De l'IA à la superintelligence | Intelligence Artificielle 25 62K views
ALPHAZERO - ORÍGENES Y FUNCIONAMIENTO - EVOLUCIÓN DE LA IA EN EL JUEGO DE AJEDREZ PARTE 2 51K views
AlphaZero - Stockfish: ТОРЖЕСТВО ДУХА НАД МАТЕРИЕЙ! 45K views
AlphaZero - Stockfish: БОРЬБА СЛОНОВ ПРОТИВ КОНЕЙ 41K views
Outrageous Chess AI: (Game 10) : DeepMind’s AlphaZero's outrageous Queen moves from other dimension! 40K views
Stockfish - AlphaZero: поучительный УРОК СТРАТЕГИИ в закрытых позициях! 38K views
Stokfish - AlphaZero: ПОЗИЦИОННАЯ ЖЕРТВА ФИГУРЫ в испанской партии 37K views
AlphaZero - Stockfish: СЛОЖНАЯ СТРАТЕГИЧЕСКАЯ БОРЬБА 35K views
Alpha Zero VS Stockfish | Legendäre Partie | Künstliche Intelligenz im Schach 31K views
Outrageous Artificial Intelligence (Game 2): DeepMind’s AlphaZero crushes Stockfish 29K views
Jan Gustafsson on his game, Google Deepmind Alpha Zero and his Opening Clinic format on chess24 26K views
Stockfish CRUSHED by Google's neural network Alpha Zero 26K views
ImageNet Moment for Reinforcement Learning? 26K views
Outrageous Chess AI: (Game 5) : Deepmind's AlphaZero: One of the most outrageous moves of the year! 22K views
Outrageous Artificial Intelligence: (Game 3) : DeepMind’s AlphaZero crushes Stockfish Chess computer 21K views
Awesome Annotations #3 - AlphaZero Hulk-Smashes Stockfish 21K views
Le phénomène AlphaZero contre Stockfish 21K views
AlphaZero vs Stockfish 8 | Jogo 10 | IA do Google faz lances de quebrar a cabeça 20K views
Inteligência Artificial do Google devasta Stockfish em 100 partidas 20K views
Outrageous Artificial Intelligence: (Game 7) : DeepMind’s AlphaZero crushes Stockfish Chess Engine 20K views
Google's Artificial Intelligence Alpha Zero Conquers Chess Only 4 Hours After Learning The Rules! 18K views
ПРЕТВОРЯЯ В ЖИЗНЬ: Живые существа, созданные человеком 17K views
AlphaZero gives up a pawn to bury another Stockfish bishop 17K views
The Evolution of AlphaGo to MuZero 15K views
AlphaZero gives Stockfish another positional lesson 15K views
AlphaZero - Stockfish: жертва пешки во французской защите 15K views
Outrageous Artificial Intelligence: (Game 4): French Defence: DeepMind’s AlphaZero crushes Stockfish 15K views
Outrageous Artificial Intelligence: (Game 6) : DeepMind’s AlphaZero crushes Stockfish Chess Engine 14K views
AlphaZero AI teaches itself chess and crushes Stockfish!! 14K views
Moving Beyond Surface Statistics (Apple researcher) 12K views
将棋界最強AlphaZeroの論文をざっくり解説 11K views
Outrageous Chess AI: (Game 9) : Classic effective strategy against Classic French defence downsides 11K views