In artificial intelligence history, the ancient board game of Go was considered the ultimate grand challenge. With more board configurations than there are atoms in the observable universe, traditional brute-force chess supercomputers like Deep Blue were utterly useless against human masters who relied on aesthetic intuition.

Google DeepMind designed an AI that learned through millions of games of self-play. By pairing an intuition network that spots beautiful moves with an evaluation network that calculates long-term win probabilities, AlphaGo played with a mysterious, creative elegance—delivering historic moves that overturned three thousand years of human Go strategy.

AlphaGo demonstrated that machines could discover new knowledge beyond human teaching. By laying the algorithmic foundations for AlphaFold to solve the 50-year protein folding challenge, by optimizing global energy grid cooling, and by accelerating modern reasoning AI, deep reinforcement learning opened the path toward Artificial General Intelligence.

Reference

Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., & Hassabis, D. (2016). Mastering the game of Go with deep neural networks and tree search. Nature, 529(7587), 484–489.

Title

Mastering the game of Go with deep neural networks and tree search

Abstract

The game of Go has long been viewed as the most challenging of classic games for artificial intelligence owing to its enormous search space and the difficulty of evaluating board positions and moves. Here we introduce a new approach to computer Go that uses 'value networks' to evaluate board positions and 'policy networks' to select moves. These deep neural networks are trained by a novel combination of supervised learning from human expert games, and reinforcement learning from games of self-play. Without any lookahead search, the neural networks play Go at the level of state-of-the-art Monte Carlo tree search programs that simulate thousands of random games of self-play. We also introduce a new search algorithm that combines Monte Carlo simulation with value and policy networks. Using this search algorithm, our program AlphaGo achieved a 99.8% winning rate against other Go programs, and defeated the human European Go champion by 5 games to 0. This is the first time that a computer program has defeated a human professional player in the full-sized game of Go, a feat previously thought to be at least a decade away.

Cited 15,929 times · View on doi.org

Continue

Continue Exploring

Ask this paper your own questions, or keep browsing the verified research catalogue.