Ataraxos study expands hidden information AI beyond Stratego
The expanded Ataraxos study tests its approach in three more games, with separate evidence for human competition and AI benchmarks.

A Nature study published September 30 extends the Ataraxos research beyond classic Stratego, reporting strong results in Barrage Stratego, Hanabi and dou dizhu. The expanded experiments test competing, cooperating and team play under hidden information.
The original Stratego victory appeared in a November 2025 preprint. That earlier paper focused on Stratego. The broader experiments in the new publication provide the substantive update, rather than a newly announced win against the same human opponent.
What the additional games show
In Barrage Stratego, Ataraxos won four series of 50 games against three highly ranked human players. The researchers describe this as the first superhuman result in that variant.
Hanabi testing covered two through five players. In dou dizhu, the opponents were PerfectDou and DouZero. The paper reports leading AI results in both card games. Those comparisons should not be confused with human championship victories.
How the system handles hidden pieces
MIT’s account of the research describes two connected stages. First, self play produces a learned strategy. During an actual game, a model estimates plausible identities for the opponent’s concealed pieces. Planning then refines the next decision before the system acts.
This approach avoids having to inspect every possible hidden arrangement. The distinction matters in Stratego, where a useful move depends on what an opponent might know, what they might be concealing and how they interpret a player’s behavior. The system combines learning before play with additional computation during play.
The earlier Stratego result needs context
The original evaluation recorded 15 wins, one loss and four draws against Pim Niemeijer. A separate demonstration at the August 2025 World Championship produced 38 wins and two losses against attendees. These were demonstration games, not a championship title.
The authors estimated training costs below $8,000 at 2025 rental prices. That included 16 H100 GPUs for a week of reinforcement learning and four H100s for four days of belief model training. Their DeepNash comparison estimated $3 million to $4.5 million using a reported hardware setup and a researcher’s recollection of training time.
These estimates concern training hardware costs. They do not establish a complete research budget or a precisely measured 500 fold reduction in computation. The authors also report that a direct match with DeepNash could not be arranged.
Code offers a route to independent testing
The team has public repositories for Stratego, Hanabi and dou dizhu. The Hanabi repository separates policy training, belief model training and search evaluation. Its configurations cover different player counts, making the stages of the approach inspectable.
The dou dizhu repository likewise includes training, inference, search and evaluation components. This is research software for reproducing and extending experiments. A public repository alone does not establish that an independent team has reproduced the published results.
Real world applications remain prospective
MIT identifies negotiations and cybersecurity as possible future applications. It also says the researchers want to make the system’s decisions understandable enough for people to audit. The demonstrated results remain game experiments. Reliable performance in a live business or security setting still needs its own evidence.
Read more studies and benchmark coverage in ByteForward’s Research section.
Featured illustration by Samuel Sokota and colleagues from their original Stratego paper, used under CC BY 4.0. No changes.



