QUEEN is a hybrid system that makes a chess engine speak and reason in natural language by coupling a strong chess network (Leela/Lc0) with a small encoder-decoder language model (SmolLM3-3B) via learned “bridge” parameters. The bridge is first trained on templated QA about static and dynamic board features, then bootstrapped with ~8.4K distilled explanations from GPT-5.6-Sol (low). The key algorithmic innovation is a natural-language analogue of AlphaZero: have an LM generate explanations for leaf nodes of a search tree, merge those explanations upward through the tree, and use the aggregated root explanations as supervised fine-tuning data for the next iteration. This creates a self-improving loop that trains the LM to both play and explain chess coherently.
Empirically, iterative self-distillation transforms the system from beginner/intermediate to roughly grandmaster strength on Lichess blitz ratings over seven iterations, outperforming other frontier LMs on Elo and puzzle solving while approaching higher-tier models in explanation quality. The authors also demonstrate a path that avoids heavy external distillation: starting from templatic explanations and letting an LM (Qwen) combine and refine them through self-improvement raised a model from 2115 to 2497 in three iterations. The approach generalizes beyond chess, combining architectural and algorithmic advances to produce interpretable, game-strong language-capable agents.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.