Chess Puzzles
100 novel puzzles, generated programmatically with a chess engine. Each puzzle has a single best next move.
100 novel puzzles, generated programmatically with a chess engine. Each puzzle has a single best next move.
The benchmark consists of chess positions where models must select the best next move. All positions were generated programmatically by Epoch and do not appear in any other source. Each position has a single best move for the player whose turn it is, as judged by the Stockfish chess engine. Models are graded based on whether they select the correct move.
While the broader significance of chess is low, we find this benchmark useful for several reasons.
Models are given the board state via Forsyth–Edwards Notation (FEN); they are not asked to process an image. You can see model solutions using our log viewer, for instance here are the logs for GPT-5. Several websites, e.g. Chess.com, have tools to load FEN strings, display them as board states, and analyze positions.
Models are given the following prompt.
Below is a FEN string representing the state of a chess game. Please analyze the position and determine the best next move for the player whose turn it is. You may think as much as you would like, but please conclude your response with the move expressed as a four-character string stating the starting and ending square of the piece to be moved, using the format "MOVE: b1c3" (without quotes).
{FEN}
A model-based answer extractor is tasked with extracting the model’s final answer. This is then compared for an exact match to the answer key.
Board states are generated by a heuristic search and filtering program. The first six full moves are played randomly. Subsequently, moves are played by Stockfish set to a randomly-drawn Elo rating. After a game concludes, full-powered Stockfish is used to find positions where there is a single best move that the player model did not select. Positions are filtered out if the best move is capturing an undefended piece or a higher-value piece with a lower-value piece as these usually constitute obvious blunders. We found that one desired position was generated roughly every five simulated games.
The implementation of the puzzle generator, as well as the Inspect task definition, can be found here. The dataset was generated in a single run, with “chaos” mode enabled. We thank Gemini 3 Pro Preview for the vibe coding.
Our subjective assessment of the positions is that they range in difficulty from somewhat straightforward to somewhat difficult. We show one example of each below.
Two positions featured in the benchmark. White’s only good move in the position on the left is relatively straightforward, even if the reasons it is the only good move are less clear. White’s only good move in the position on the right is far less obvious.
To put model scores in context, we evaluated several human players on a sample of puzzles from the benchmark.
Five players each answered a sample of 20 randomly chosen puzzles. Players were asked to work alone, without use of AI or a chess engine. Each player was asked to identify the single best move in each position and to record the time taken to answer.
Chess ability varied across players, with US Chess Federation (USCF) ratings ranging from 1700 to 2400 (one player instead reported a rating of 2000 from a popular online chess website). One player was accidentally shown a chess engine’s solution to one puzzle; that puzzle was excluded when calculating their score.
Results
Strong human players generally outperformed AI models on the evaluated puzzles. The two highest-scoring players both answered 75% of puzzles correctly, while the highest score among models on the same puzzle set was 60% (GPT-5.5 Pro at xhigh reasoning effort, and Gemini 3.8 Flash). A majority of models scored 20% or lower. GPT-6 Astra (max) scored slightly lower on the sample (55%) despite leading the benchmark overall (72%).
That the top human score was 75% shows that these puzzles are very challenging for humans, at least when spending on the order of minutes per puzzle. The sampled puzzles are also harder than the full puzzle set on average: the mean AI score is 8 percentage points lower on the sample than on the full set.
Two positions were not solved correctly by any human player. One of these was solved by three models (GPT-5.5 Pro, GPT-5.2 and Gemini 3.8 Flash), while the other was solved only by GPT-6 Astra.
Anonymized results for both human players and models are available in this spreadsheet.
Out of 19 puzzles; one puzzle was excluded after this player was accidentally shown an engine solution.
This player provided an online (Chess.com) rating rather than an over-the-board rating.
100 novel puzzles, generated programmatically with a chess engine. Each puzzle has a single best next move.