Back to benchmarks

Lech Mazur Writing

A creative story-writing benchmark where models must weave ten assigned elements into short stories, graded on a 0–10 rubric by a panel of LLM judges. Frozen at the benchmark's final 0–10 leaderboard (August 2025).

Lech Mazur Writing

Update: Our data for this benchmark is frozen at the final leaderboard of its original version, published August 8, 2025, which scored 50 models on an absolute 0–10 scale. The benchmark has since moved to a new scoring system based on pairwise story comparisons, which produces relative ratings that cannot be combined with the 0–10 scores shown here (see Methodology). This page is no longer updated.

The Lech Mazur Writing benchmark (the lechmazur/writing project) tests how well models write short creative stories under rigid constraints. Each model writes 500 stories of roughly 400–500 words, and every story must organically incorporate ten assigned random elements: a character, object, core concept, attribute, action, method, setting, timeframe, motivation, and tone. A panel of seven LLM graders scores each story on a 16-question rubric covering character development, plot structure, atmosphere, storytelling craft, originality, cohesion, and how well each of the ten required elements is integrated. The reported score is the mean rating on a 0–10 scale.

Because every story shares the same required building blocks and similar length, outputs are directly comparable across models, and the benchmark captures both constraint satisfaction and literary quality rather than fluency alone.

Methodology

Scores are sourced from the final version 1 leaderboard of the lechmazur/writing repository (snapshot of August 8, 2025), the last release to grade stories on an absolute 0–10 scale.

Later versions of the benchmark no longer assign absolute grades. Instead, evaluator models compare pairs of stories head-to-head, and the pairwise judgments are aggregated into Thurstonian comparison ratings. These ratings are zero-centered and scaled relative to whichever models are in the comparison pool, and they are re-estimated as models are added or removed — a model’s rating can change without its stories changing. Ratings of this kind cannot be used in the Epoch Capabilities Index, which requires benchmark scores that map onto a fixed 0–1 scale: pool-relative ratings have no stable correspondence to an absolute level of performance, and mixing them with the earlier 0–10 grades would conflate changes in the comparison pool with changes in model capability. We therefore freeze this benchmark at its final 0–10 leaderboard and treat it as a completed, historical benchmark.