A reinforcement learning model that learns board-game strategy by playing itself.
AlphaZero is the kid who plays both sides of the chessboard at lunch. It beats itself, gloats a little, then comes back stronger.
It starts with the rules, then learns by self-play. It is famous for Go, chess, and shogi.
AlphaGo
AlphaZero is a more general AlphaGo, built for more than Go.
RL
AlphaZero improves by playing itself and learning from wins and losses.
MCTS
AlphaZero uses MCTS to search ahead and pick stronger moves.
Deep RL
AlphaZero is one famous example of Deep RL.