Queen: a 4-billion-parameter model that plays like a Grandmaster and explains its moves
This paper introduces Queen, a chess-focused language model that both plays strong chess and explains its choices in natural language. The authors build on two observations: modern chess engines play extremely well but do not explain their moves, while language models can explain moves but often play weakly. Queen aims to combine the strengths of both by pairing a high-quality chess “encoder” with an instruction-tuned language “decoder.”
The model uses an encoder-decoder design inspired by prior multimodal work. The encoder is Lc0, a strong transformer-based chess network that produces internal position representations. The decoder is an instruction-tuned language model that receives those representations through a cross-attention mechanism. The team trains only the cross-attention and related components with a question-and-answer curriculum that moves from understanding static boards to predicting how moves change positions.
On top of that architecture, the researchers introduce an iterative distillation method they describe as a natural-language analogue of the Bellman update (a standard idea in sequential decision-making). In each round the model analyzes the positions that follow its top candidate moves, combines those child analyses into a consolidated explanation of the current position, and then distills that explanation back into the model. After seven such iterations the model’s estimated playing rating rose by over 900 Elo points, from 1782 to 2697.
These gains make Queen stronger than several large frontier language models on the authors’ tests. The paper reports that Queen’s estimated rating (2697) exceeds GPT-5.6-Sol (2071) and Gemini-3.1-Pro (2201), and approaches the median Grandmaster Lichess blitz rating (2730). The authors also measure explanation quality along three axes: accuracy (playing strength), substantiation (how well the model’s predicted move consequences match tactics), and coherence (natural-language fluency). On tactical puzzles Queen had a higher no-mistake rate than baselines, and LM-based judges found its explanations fluent and nearly as coherent as GPT-5.6-Sol.