hn.today

With most information hidden, the game Stratego had stumped AI–until now

arstechnica.com7 points0 comments
Screenshot of With most information hidden, the game Stratego had stumped AI–until now

A new AI called Ataraxos has finally cracked Stratego, a long-elusive imperfect-information game, by beating Pim Niemeijer - the game’s top-ranked player - 15 games to one with four draws in a 20-game match, and winning 38 of 40 exhibition games at the 2025 World Championship. Stratego’s challenge comes from 40 hidden pieces per side, decillion-plus setup possibilities, long games that can run thousands of moves, and pervasive bluffing. Ataraxos pairs conventional self-play reinforcement learning (about 163 million games) with two crucial twists: an explicit belief network that predicts opponents’ hidden piece identities from their moves, and a search step that samples plausible hidden setups and evaluates candidate moves across them. Training used an aggressive early update schedule that tapers into finer adjustments later, which stabilizes learning in the face of massive hidden information.

The system is strikingly efficient: training ran on 16 GPUs for a week plus four GPUs for four days to train the belief model, a fraction of the compute and cost DeepNash required. Ataraxos plays calmly and methodically, intentionally making human-like bluffs and avoiding moves that would reveal hidden weaknesses, and its unusual plays (for example tucking a flag behind two bombs) have already shifted human metagame patterns. The architecture also generalized to Barrage Stratego, Hanabi, and dou dizhu, and researchers are targeting broader applications in modeling adversarial real-world problems while working to make strategies more interpretable.

Read on arstechnica.com0 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.