Researchers build a superhuman Stratego agent by combining self-play reinforcement learning with a novel test-time search method tailored to games with massive hidden information. Stratego has long resisted AI breakthroughs that rival top human play because of imperfect information and the size of the strategy space. The team trains agents via self-play to learn strategic priors and then applies an imperfect-information-aware search at decision time to refine choices, producing a large improvement in playing strength versus previous efforts and top humans.
The approach emphasizes general techniques for self-play under uncertainty and search mechanisms that operate effectively when opponents’ piece identities are unknown. Importantly, the implementation achieves vastly superhuman performance while keeping training and evaluation costs low - reported to require only a few thousand dollars rather than the historically cited million-dollar budgets for similar milestones. The work demonstrates both a methodological advance for practical imperfect-information game AI and a cost-performance breakthrough, with implications for other domains requiring strategic decision-making under hidden information.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.