MIT-Led Team's Ataraxos Beats the Strongest Human at Stratego, Using Less Than One Hundredth of DeepNash's Training Examples

1 views

On September 30 a research team from MIT, Carnegie Mellon University, New York University and Stanford University published a Nature paper, "Scalable decision-making for games of imperfect information," introducing Ataraxos. It reaches superhuman level at Stratego, a board game in which each side's piece identities are hidden from the opponent: over 20 games against Pim Niemeijer, described as the most decorated Stratego player of all time, it went 15 wins, 1 loss and 4 draws, and MIT says it went 39-2 against top players at the Stratego world championship. Each side secretly deploys 40 pieces, and there are more than 10 to the 66th power possible setups. Ataraxos first learns a blueprint strategy through self-play reinforcement learning, then plans at decision time during a game: a generative model estimates the most likely identities of the opponent's hidden pieces, samples possible board states and simulates candidate moves before moving. The researchers say it plays strictly stronger than DeepMind's 2022 DeepNash while using less than one hundredth of its training examples and less than one thirtieth of its self-play games. The same approach also reached superhuman or state-of-the-art play in Barrage Stratego, the cooperative card game Hanabi and the Chinese card game dou dizhu.

Why Stratego is harder than Go

Go and chess are perfect-information games: both players see the same board. Stratego is a close relative of the Chinese game junqi: you know where the opponent's pieces are but not whether a piece is a marshal or a bomb, and identities are revealed only when pieces fight. Players have to guess while deliberately getting the opponent to guess wrong.

Problems like this used to be tackled mainly with methods from poker AI, but Stratego's hidden state is far larger than poker's. Gabriele Farina of MIT, the paper's senior author, says techniques developed for poker "definitely could not scale in this setting." DeepMind's DeepNash reached near top-human level in 2022 with massive compute but could not consistently beat the strongest players.

Being cheap is the real point of the paper

Ataraxos's breakthrough isn't only that it wins, but that it wins with far less training. Rather than working everything out during training, it plans before each move: it first guesses what the opponent most likely holds, then reasons forward under those assumptions. In Farina's words, rather than guessing blindly, it uses decision-time planning to find the most plausible state of the board.

For AI research, the point is that "learn more in training, compute less at inference" is too expensive, and shifting some computation to decision time can solve imperfect-information problems on a much smaller training budget. It mirrors the idea in large language models of spending compute at inference time.

How far it is from real use

The team names negotiation and cybersecurity, settings where decisions are made with incomplete information, as potential uses. They also acknowledge the system currently has trouble explaining why it moves the way it does; the next step is interpretability work so people can audit its recommendations before any real-world use.

Scope note: MIT reports a 39-2 record "at the world championship"; some coverage notes these were demonstration games against attendees during the championship, not entry into and victory at the official human tournament. The team did not play Ataraxos against DeepNash directly; the efficiency comparison is based on each system's published training scale.

via: MIT News, Nature paper, TechXplore report