Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Ataraxos study expands hidden information AI beyond Stratego

The expanded Ataraxos study tests its approach in three more games, with separate evidence for human competition and AI benchmarks.

Listen to this article

A Nature study published September 30 extends the Ataraxos research beyond classic Stratego, reporting strong results in Barrage Stratego, Hanabi and dou dizhu. The expanded experiments test competing, cooperating and team play under hidden information.

The original Stratego victory appeared in a November 2025 preprint. That earlier paper focused on Stratego. The broader experiments in the new publication provide the substantive update, rather than a newly announced win against the same human opponent.

What the additional games show

In Barrage Stratego, Ataraxos won four series of 50 games against three highly ranked human players. The researchers describe this as the first superhuman result in that variant.

Hanabi testing covered two through five players. In dou dizhu, the opponents were PerfectDou and DouZero. The paper reports leading AI results in both card games. Those comparisons should not be confused with human championship victories.

How the system handles hidden pieces

MIT’s account of the research describes two connected stages. First, self play produces a learned strategy. During an actual game, a model estimates plausible identities for the opponent’s concealed pieces. Planning then refines the next decision before the system acts.

This approach avoids having to inspect every possible hidden arrangement. The distinction matters in Stratego, where a useful move depends on what an opponent might know, what they might be concealing and how they interpret a player’s behavior. The system combines learning before play with additional computation during play.

The earlier Stratego result needs context

The original evaluation recorded 15 wins, one loss and four draws against Pim Niemeijer. A separate demonstration at the August 2025 World Championship produced 38 wins and two losses against attendees. These were demonstration games, not a championship title.

The authors estimated training costs below $8,000 at 2025 rental prices. That included 16 H100 GPUs for a week of reinforcement learning and four H100s for four days of belief model training. Their DeepNash comparison estimated $3 million to $4.5 million using a reported hardware setup and a researcher’s recollection of training time.

These estimates concern training hardware costs. They do not establish a complete research budget or a precisely measured 500 fold reduction in computation. The authors also report that a direct match with DeepNash could not be arranged.

Code offers a route to independent testing

The team has public repositories for Stratego, Hanabi and dou dizhu. The Hanabi repository separates policy training, belief model training and search evaluation. Its configurations cover different player counts, making the stages of the approach inspectable.

The dou dizhu repository likewise includes training, inference, search and evaluation components. This is research software for reproducing and extending experiments. A public repository alone does not establish that an independent team has reproduced the published results.

Real world applications remain prospective

MIT identifies negotiations and cybersecurity as possible future applications. It also says the researchers want to make the system’s decisions understandable enough for people to audit. The demonstrated results remain game experiments. Reliable performance in a live business or security setting still needs its own evidence.

Read more studies and benchmark coverage in ByteForward’s Research section.

Featured illustration by Samuel Sokota and colleagues from their original Stratego paper, used under CC BY 4.0. No changes.

Jordan Reid
Jordan Reid

Jordan Reid is focused on AI tools, agents, developer products, and the way technology changes everyday work. Jordan approaches a launch from the user’s side of the screen. What can it actually help someone finish? The voice is practical, conversational, and skeptical of products that turn a simple job into five new settings. Coverage follows coding assistants, creative software, browser agents, and the workflows around them, with attention to pricing, permissions, setup, and the human work that remains.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile