Brood War Bench is a round-robin benchmark that pits 19 model/effort configurations against one another in StarCraft: Brood War using an automated harness running parallel matches on Freestyle VMs. Matches produced a leaderboard (wins, losses, APM, cost per game), time-series metrics (workers, army size, structures, unspent resources, tech/upgrades), and a full head‑to‑head matrix. Every configuration played every other, games were logged with both agent harness output and game-engine data, and finished-game samples drop out of the time series rather than being filled.
Results show none of the agents played beyond a beginner level, but Codex Astra clearly led the field - Astra xhigh went 18-0 (100% win rate) and Astra medium 16-2. Claude Fable was the most earnest at executing multi-stage strategies, sometimes reaching Lair/Spire/Mutalisks or higher tech before winning, while Grok 4.6 frequently stalled in long stretches of reasoning and issued almost no commands (examples include runs with thousands of reasoning tokens but only a handful of command batches or no combat units fielded). Common failure modes included cheap disruption/cheese (probe harassment), poor coordination among subagents that sent units one-by-one, and older models treating the RTS like a turn-based game and losing while thinking. Most games ended in the first 5-15 minutes (Astra median 8:10, Fable 10:37, Grok 9:15).
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.