Ataraxos Defeats Stratego’s Best Human Player With Hidden-Information Reasoning
Summary
Stratego had remained a difficult target for AI because players must reason about 40 hidden pieces across games that can last up to 2,000 moves, while also bluffing. Researchers from Carnegie Mellon, MIT, NYU, and Stanford developed Ataraxos, which beat four-time world champion Pim Niemeijer 15 games to one, with four draws, in a 20-game online match; Niemeijer earned $100 per win. Ataraxos trained through 163 million self-play games, using larger strategy updates early in training and smaller updates later to reduce instability caused by hidden information. Its key addition was a second neural network, called a belief model, that inferred likely identities of hidden pieces from the opponent’s moves. Before each action, the system sampled plausible board arrangements, simulated candidate moves, and selected among the outcomes instead of enumerating every possible setup. The team trained the main system on 16 GPUs for a week and the belief model on four GPUs for four days, at a cost of a few thousand dollars, while the researchers estimated DeepMind’s DeepNash required $3 million to $4.5 million in specialized-chip costs. Ataraxos also won 38 of 40 games against visitors at the 2025 Stratego World Championship, beat three world champions in Barrage Stratego, mastered Hanabi, and outperformed leading bots in dou dizhu. The researchers say the approach may apply to simplified models of negotiations, markets, or conflict, but Ataraxos currently cannot explain why it chooses its moves and remains less interpretable than they want.