AI defeats Stratego champion by exploiting clues from hidden pieces
A research team from Carnegie Mellon, MIT, New York University and Stanford has built Ataraxos, an AI that defeated leading Stratego player Pim Niemeijer 15 games to one, with four draws. The result matters because Stratego’s hidden pieces, bluffing and long games have made it unusually difficult for computers to master, despite earlier successes in chess, Go and poker.
Ataraxos trained through 163 million self-play games using 16 GPUs and a budget of a few thousand dollars. Its key addition is a second neural network that estimates the opponent’s hidden pieces from their moves, allowing the AI to sample plausible board positions and search ahead without evaluating every possible setup. The researchers also adjusted its learning strategy over time, making larger changes early in training and smaller ones later.
- Ataraxos beat Pim Niemeijer 15–1, with four draws.
- It trained on 16 GPUs for a few thousand dollars.
- A network for guessing hidden pieces helped it plan ahead.
New here? Start with this
Stratego is a two-player board game where armies battle to capture each other's flags. Each player's pieces start hidden from their opponent, and they must guess where enemy troops are positioned as the game unfolds. The hidden information and bluffing element make Stratego fundamentally different from games like chess.
Artificial intelligence has already mastered chess and Go, games requiring strategic thinking but where both players see all pieces at all times. Stratego has proven far more resistant because computers must estimate what hidden pieces might be, leading to an enormous number of possible board states to consider. This has made Stratego one of the most challenging competitive board games for artificial intelligence.
Researchers from Carnegie Mellon, MIT, New York University and Stanford developed Ataraxos to tackle this challenge. Their key innovation was a neural network trained to estimate where hidden pieces are likely positioned based on an opponent's moves, letting the AI sample probable board states rather than evaluate all possibilities. The system was trained through playing millions of games against itself.
Read the full article at the source →
Originally published by Ars Technica as “With most information hidden, the game Stratego”.