IA4 MIN

Ataraxos takes its hidden-information AI beyond Stratego

A new Nature paper extends the approach to Barrage Stratego, Hanabi and dou dizhu. The researchers train separate systems using a shared method for acting on incomplete information.

Stratego boards set up at the 2024 Patras Battles tournament, an archive photograph unrelated to the Ataraxos tests
Image: Catlemur / Wikimedia Commons · CC BY-SA 4.0 · Patras Battles 2024 · resized / redimensionada

The method behind Ataraxos is working beyond the Stratego board. A Nature paper published September 30 reports wins against leading human players in Barrage Stratego and stronger results than previous AI systems in Hanabi and dou dizhu. Each game uses a separately trained system built around the same approach to hidden information.

The original Stratego victory was already described in a November 2025 preprint. The broader study tests whether the approach also works when players need to cooperate. Hanabi puts everyone on one team, while dou dizhu pits a pair of players against a third. This is evidence for a reusable training method, rather than one model learning a game and immediately mastering the others.

01

Planning around pieces it cannot identify

Stratego players secretly arrange 40 pieces each. An opponent can see where a piece stands without knowing whether it is weak, powerful or a bomb. Combat reveals identities. Until then, a useful move has to account for several possible boards without giving away too much about your own pieces.

Ataraxos develops a strategy by playing against copies of itself and gradually favoring decisions that lead to wins. This is self-play reinforcement learning. The researchers regulate how strongly the strategy changes during training so that a new adjustment does not wipe out earlier progress.

A second component learns to estimate plausible identities for the hidden pieces. Before making a move, Ataraxos tries candidate actions across several possible versions of the board, assesses the positions they could lead to, then refines its choice. It can weigh risks without identifying every enemy piece correctly. Learning to win is a different job from the dialogue generation and graphics reconstruction covered in our guide to AI in video games.

Blue Stratego pieces with identities visible from their owner’s side, in an illustrative photograph from 2015
Image: AWFly / Wikimedia Commons · CC BY-SA 4.0 · archivo / archive (2015) · resized / redimensionada
02

What the expanded tests actually show

In the original Stratego evaluation, Ataraxos beat four-time world champion Pim Niemeijer with 15 wins, one loss and four draws across 20 games. Niemeijer could study the system between games. Ataraxos did not retrain to adapt to him.

Barrage Stratego reduces each player’s army from 40 pieces to eight. The researchers tested their systems in four 50-game series against three two-time world champions and won every series. Three series used the learned strategy without additional search before a move. The fourth included that search.

Hanabi reverses the information problem. Players can see their teammates’ cards but not their own, and must jointly build five color-coded sequences in order. In five-player testing, Ataraxos achieved a perfect score in about 58% of games, compared with the previous best result of 32.9% cited by the paper. The researchers evaluated each condition over 10,000 games.

Dou dizhu requires players to empty their hands, with two participants cooperating against a third. Ataraxos beat the existing PerfectDou and DouZero programs. The card-game results compare AI systems, so they should not be described as victories over human champions.

Numbered cards in several colors and tokens from Hanabi, in an illustrative photograph unrelated to the Ataraxos tests
Image: Mannivu / Wikimedia Commons · CC BY-SA 4.0 · fotografía ilustrativa / illustrative photo (2020) · resized / redimensionada
03

What training requires beyond the hardware

The authors estimate that training the full Stratego system would cost less than $8,000 at 2025 hardware rental prices. That is a compute estimate, not the total research budget. Learning to play used 16 NVIDIA H100 accelerators for a week. Training the component that estimates hidden pieces used four for four days.

A fast simulator helped make that possible by processing millions of board-state changes per second. Useful practice depends on an environment that responds quickly and follows the rules correctly. The need for a reliable practice environment also appears in Qwen-AgentWorld, which simulates computer tools rather than Stratego matches.

Applications such as negotiations or cybersecurity remain research possibilities. Success in a board game does not establish success in a setting where accurate simulation is harder. In MIT’s account of the work, the researchers also identify a further goal: making Ataraxos’s decisions understandable enough for people to inspect before following its recommendations.

00

The conversation starts here

Sign in with a supporter account to comment. Sign in

Nobody has commented yet. Want to go first?

KEEP READING

You may also like

FRONT PAGE