AI Reaches Superhuman Stratego Performance with Self-Play and Test-Time Search
Summary
A new paper presents an AI system for Stratego, a board wargame built around strategic decisions with large amounts of hidden information. The authors say earlier efforts requiring millions of dollars in training costs had failed to match top human performance. Their system combines self-play reinforcement learning with test-time search methods designed for imperfect-information settings. According to the paper, the resulting agent not only reaches the level of top human players but performs vastly beyond it. The authors also report that achieving this result required only a few thousand dollars, positioning the work as both a performance and cost improvement for AI in a difficult classical game.