Researchers from Carnegie Mellon University, MIT, New York University, and Stanford University developed an AI system called Ataraxos that defeated top human players in the imperfect-information game Stratego. The system achieved a 15-1-4 victory margin against Pim Niemeijer, the most decorated Stratego player of all time.
The research team included Samuel Sokota, a graduate student at Carnegie Mellon University and lead author of the study; Eugene Vinitsky and Zico Kolter, graduate students at New York University; Hengyuan Hu, a graduate student at Stanford University; and Zhiyuan Fan, a graduate student in the Department of Electrical Engineering and Computer Science at MIT. Gabriele Farina, an assistant professor in the Department of Electrical Engineering and Computer Science at MIT and a principal investigator at the Laboratory for Information and Decision Systems, served as a senior author on the project published in Nature. "You only have two hidden cards," Gabriele Farina said.
"In Stratego, there’s 40 pieces on the board that could be in any order," Farina said. "In the kind of imperfect information tasks you would face in reality, you often don’t have the luxury of enumerating through all the possibilities," he said.
Stratego is a two-player game where each player has 40 pieces representing military ranks, bombs, and a flag, with the objective being to capture the opponent’s flag. It is an imperfect-information game because piece identities are hidden and only revealed when two pieces collide, resulting in the removal of the weaker piece. There are more than 10 to the 66th power possible piece configurations in Stratego.
"There’s something super distinctive about Stratego, which is that it is a massive amount of hidden information that unfolds over a very long time scale," said Eugene Vinitsky, a researcher at NYU. "In chess, usually the game lasts 40 moves, but in Stratego, a game can easily last 2,000 moves," he said.
Ataraxos was trained using self-play reinforcement learning to create a blueprint strategy. The system uses decision-time planning with a generative model to estimate opponent piece identities and evaluate future choices during gameplay. Training Ataraxos required 16 GPUs and cost a few thousand dollars. Specifically, the reinforcement learning training run for Ataraxos used 16 NVIDIA H100 GPUs for 1 week and 4 H100s for 4 days, costing less than US$8,000 at 2025 prices.
DeepMind introduced an AI system called DeepNash in 2022. DeepNash was trained on 1,024 tensor processing unit nodes and would have cost between US$3,000,000 and US$4,500,000 under 2025 pricing. Ataraxos uses less than one hundredth of the training examples and less than one thirtieth of the self-play games compared to DeepMind's DeepNash. Ataraxos achieved a 39-2 record against top human players at the Stratego world championship.
Why It Matters
The development of Ataraxos represents a shift in artificial intelligence capabilities for imperfect-information environments. Previous AI milestones, such as Deep Blue defeating Garry Kasparov at chess in 1997 and AlphaGo defeating Lee Sedol at Go in 2016, occurred in perfect-information settings where all game states are visible. Stratego presents unique challenges due to its hidden information and long time scales, making previous AI techniques insufficient. The ability of Ataraxos to achieve high performance with a fraction of the computational cost of prior systems like DeepNash suggests that general-purpose AI algorithms can now handle complex, real-world tasks where enumerating all possibilities is not feasible.
Timeline
Deep Blue defeated Garry Kasparov at chess in 1997. AlphaGo defeated Lee Sedol at Go in 2016.
What's New
Later reporting showed that "with Stratego, there is an explosion of possible universes you might have to deal with." Additional facts established that there are more than 10 to the 66th power possible piece configurations in Stratego. The research team that developed Ataraxos included Samuel Sokota (Carnegie Mellon University), Eugene Vinitsky and Zico Kolter (New York University), Hengyuan Hu (Stanford University), and Zhiyuan Fan (MIT). Further reporting quoted Farina stating that "in the kind of imperfect information tasks you would face in reality, you often don’t have the luxury of enumerating through all the possibilities." Additional statements included Farina saying that "our system reaches strictly higher playing strength than DeepNash (DeepMind’s system) while using less than one hundredth of the training examples and less than one thirtieth of the self-play games, indicating a massive improvement in efficiency." He also stated that "AI techniques that were developed for games like poker definitely could not scale in this setting." Later accounts confirmed that Pim Niemeijer, the most decorated Stratego player of all time, was defeated by Ataraxos in a 20-game series with a margin of 15 wins, 1 loss, and 4 draws. Additional reporting noted that Ataraxos achieved an 85% effective win rate against top human players in Stratego, which is unprecedentedly large for the highest level of human play.
forum Comments (0)
No comments yet. Be the first to comment.