Lech Mazur Writing tests a model’s ability to generate short stories that incorporate mandatory story elements. Since the adoption of pairwise scoring in April 2026, the benchmark has not maintained proper versioning. The current leaderboard combines results from different evaluator panels, bridged by an undisclosed procedure, and its scores and rankings have changed over time without documentation of the changes.
Methodology
We inspected past commits to the benchmark’s GitHub repository to track these unversioned changes. Prompts and generated stories were viewable (for a subset of rated models), but the scoring methodology was only described, so it cannot be independently verified.
Representative Errors
- Unversioned evaluator roster change within the current leaderboard: The current methodology notes that when the evaluator panel changes, past and new judgments are bridged together. The current leaderboard contains results from multiple evaluator configurations (evaluator-v2 and evaluator-v3).1
- Prompt reviewability/consistency (unknown impact on current leaderboard): Prior version changes updated the output target from a 400-500 word story to a 600-800 word story. The current published prompt files contain unfilled {min_count}/{max_count} placeholders,2 and the archived version’s README discloses that word-count instructions were separately adjusted per model to hit the target.
- Coverage (potential bias in current leaderboard): Only completed stories were compared. The README notes that Claude Opus 4.7 completed 347 of 400 stories, Claude Opus 4.8 high completed 399 of 400 stories, and Claude Fable 5 high completed 395 of 400 stories.3 It is unclear what leads to stories not being completed.
The switch to pairwise grading was not done under a new version number and subsequent changes have been unversioned.4 This includes changes in late April where the leaderboard was “refreshed” with undisclosed methodology that resulted in rescaled scores, a new evaluator panel, and a different set of models (some added, some dropped). This results in non-reproducible rankings over time. On April 18, 2026,5 Claude Sonnet 4.6 Thinking 16K trailed Claude Opus 4.6 Thinking 16K with scores of 2.54 and 3.15, respectively; on April 28, 2026,6 the ordering had flipped, with scores of 2.0 and 1.56, respectively.
Rubric
1. Reviewability
| Level | Meaning | Status |
|---|---|---|
| Full | All tasks and scoring logic inspectable*, and harness/API settings used for each model (reasoning effort, token/time limits, tool access, system prompts) are fully disclosed | [proceed to 2] |
| Partial | A representative sample of tasks and scoring logic inspectable, with full harness/API settings disclosed | [proceed to 2] |
| Inadequate | Limited, biased, or no inspectable tasks or scoring logic, harness/API settings undisclosed | [stop → NEI] |
* Either publicly or privately to reviewers.
2. Scoring
This is the minimum standard required to be Verified; failing any of these items results in a Flawed verdict.
| Scoring | |
|---|---|
| Examples |
|
| Default threshold for Flawed | ≥20% of inspected sample† contains errors or there is an issue that corrupts grading at scale |
| Status |
Pass
Flag [stop → Flawed]
Not reviewed |
| Notes | — |
| Benchmark Consistency | |
|---|---|
| Examples |
|
| Default threshold for Flawed | Leaderboard has incomparable results from different versions |
| Status |
Pass
Flag [stop → Flawed]
Not reviewed |
| Notes | — |
| Elicitation | |
|---|---|
| Examples | Model elicitation is extremely constraining and is not the focus of the benchmark:
|
| Default threshold for Flawed | Substantially reduced performance compared to reasonable alternatives for the tasks |
| Status |
Pass
Flag [stop → Flawed]
Not reviewed |
| Notes | — |
| Bias in Evaluation Setup | |
|---|---|
| Examples |
|
| Default threshold for Flawed | Material model-specific advantage found |
| Status |
Pass
Flag [stop → Flawed]
Not reviewed |
| Notes | — |
† For benchmarks with >50 available tasks, we will sample a random set of 50 (stratified by category, when present). If the observed error rate is 15–25%, we will expand the sample to 100. For benchmarks with ≤50 available tasks, all available tasks will be assessed.
3. Evaluation Quality
This is the standard we would like all benchmarks to meet, but it is not necessarily disqualifying to omit or fail these items.
| Question | Status | Notes |
|---|---|---|
| Elicitation and resource adequacy: Are the resources given to models (reasoning token/turn budget, tool access, etc.) sufficient for them to perform near their ceiling? |
Sufficient
Constraining
Unreasonably constraining
Unknown
Not reviewed | — |
| Scaffold fairness: What scaffold does the leaderboard report? |
Shared common scaffold
Mix of model-specific and common scaffolds
Model-specific scaffolds
Not reviewed | — |
| Is there evidence/risk of contamination? |
Not reviewed | — |
| Has human completability been assessed? |
All tasks
Representative set of tasks
Poor implementation (Unrepresentative set of tasks, unreasonable set of participants)
Not established
Not reviewed | — |
| Score range (if possible to estimate) | — floor, — ceiling
Not reviewed | — |
| Statistical adequacy: how many runs/model (≥ 5 recommended for error bars) | — runs/model
Unknown
Not reviewed | — |
| Construct Validity |
Measures stated capabilities
Partially measures stated capabilities
Does not measure stated capabilities
Not reviewed | — |
Disclaimer
This is a partial review. We stopped inspecting once we discovered errors that exceeded our published threshold for a flawed benchmark.
We attempted to reach out to the benchmark developer prior to launch, and if they write a response to this review we will link it here. If you’re the benchmark developer, please email us at reviews@epoch.ai with a link to your response if you would like us to include it here.
-
The README’s “Archived Absolute Ratings” section notes the benchmark switched from absolute 0-10 grading to direct story comparisons, with old and new results explicitly described as not combinable. The old results are not on the current leaderboard.
-
https://github.com/lechmazur/writing/blob/231d925888/prompts_wc/prompt_wc_0.txt
-
https://github.com/lechmazur/writing/blob/231d925888/README.md#coverage-note
-
Prior to April 2026, when updates were made, models were re-evaluated and the prior leaderboard was archived. Some of these changes were versioned.
About Benchmark Reviews
Epoch AI’s Benchmark Reviews are independent reviews of external AI benchmarks. Our documentation describes the rubric behind this verdict, how we choose which benchmarks to review, and answers frequently asked questions.