1 of 11

AlphaZero's Groundbreaking Chess Strategies and the Promise of AI

Alexey Butyrev, Mngr., Sr. Data Scientist

Phoenix Data Science Meetup, May 2019

2 of 11

Presentation Plan

  • Chess Evolution
  • AlphaZero principles
  • Training AlphaZero
  • Chess Basics
  • StockFish metric
  • AlphaZero metric
  • StockFish vs AlphaZero
  • How does AlphaZero play?

3 of 11

Chess Evolution

280-550

Chaturanga (Sanskrit: चतुरङ्ग; caturaṅga)

1200-1700

Chess in Europe.

Official Rules

1700–1873

The Romantic Era in chess

1873–1945

Birth of a sport

1948–1972

Soviet Dominance

1972-1975

Bobby Fisher WCC

1997

Deep Blue beats Kasparov

2017

AlphaZero beats Stockfish

(28 wins, 0 losses, and 72 draws)

Feb 2019

High level explanation of AlphaZero

4 of 11

AlphaZero Principles�Deepmind’s approach:

  1. Learning without being programmed
  2. General rather than specific
  3. Grounded rather than logic
  4. Active rather than passive

5 of 11

Training AlphaZero:

  • During 9 hours AlphaZero played 44M games against itself (> 1,000 games per second)
  • Played against itself with no chess knowledge
  • 5,000 1st-generation TPUs to generate self play games
  • 16 2d-generation TPU were used to train NNs
  • 4 1st-generation TPU to play against Stockfish

6 of 11

Chess Basics

Opening

Middlegame

Endgame

7 of 11

Stockfish Metric

Resulting value is computed by combining Middlegame evaluation and Endgame evaluation

8 of 11

AlphaZero Metric

  1. How plausible the move in this type of position
  2. How promising is the outcome of the variation
  3. How often variation has been considered in the search

9 of 11

Stockfish Vs AlphaZero

Stockfish

AlphaZero

Designed by chess professionals

No knowledge except game rules

Hasn’t been taught any positional ideas or strategy

Uses openings data bases

No openings (invents its own openings)

Endgame tables

Knows nothing about endgames

Uses function of material as metric

Evaluates probabilistically based on chance of wining

Trying to find the best line

Chooses position that average positions after are the best

Evaluation is midgame + endgame material

Evaluation is a flexible structure

10 of 11

How does AlphaZero play?

1. AlphaZero likes to target the opponent’s king.

2. AlphaZero likes to keep its own king out of danger.

3. AlphaZero makes sure the central situation is stable before it weakens its own kingside structure to open lines against the opponent’s king.

4. Before AlphaZero launches a wing attack, it always ensures that either it controls the centre, or the centre is stable or fixed so that its opponent cannot initiate counterplay there.

5. AlphaZero is not afraid to sacrifice material (normally one or two pawns) at an early stage to open lines or diagonals against the opponent’s king.

6. AlphaZero looks to combine an open file and an open diagonal against the opponent’s king.

7. AlphaZero loves attacking with opposite-coloured bishops.

8. AlphaZero finds great outposts for its knights and is not afraid to sacrifice material to gain time to transfer them there.

9. AlphaZero excels in building up wing attacks against the enemy king with a fixed centre.

10. AlphaZero often wins games by making some of its opponent’s pieces passive, and then exchanging off the opponent’s active pieces.

11. AlphaZero defends by creating confusion and introducing tactics into the game.

12. AlphaZero is not afraid to delay occupying open files with its rooks if it thinks it can open a file against the enemy king elsewhere on the board.

13. AlphaZero looks to restrict the mobility and freedom of the opponent’s king. It exploits this factor in its plans both in the middlegame and in the endgame.

14. AlphaZero keeps an eye out for the possibility to switch to a kingside assault. If its opponent’s pieces lose coordination, it will look to start kingside operations as quickly as possible.

11 of 11

Thank you!