We use the following rubric to evaluate each benchmark.
Rubric v1. We expect future revisions to account for new types of benchmark flaws.
| Verdict | Definition |
|---|---|
| Verified | Reviewability (Section 1) was judged as Full or Partial, and no score-affecting defect (Section 2) was severe enough to cross our quality threshold. All noted issues are logged as caveats in the review's Limitations section. |
| Flawed | Reviewability (Section 1) was judged as Full or Partial, but at least one score-affecting defect (Section 2) was severe enough to cross our quality threshold. |
| Not enough info | Required information was unavailable to review, so a Verified/Flawed label cannot be supported. |
| Level | Meaning | Status |
|---|---|---|
| Full | All tasks and scoring logic inspectable*, and harness/API settings used for each model (reasoning effort, token/time limits, tool access, system prompts) are fully disclosed | [proceed to 2] |
| Partial | A representative sample of tasks and scoring logic inspectable, with full harness/API settings disclosed | [proceed to 2] |
| Inadequate | Limited, biased, or no inspectable tasks or scoring logic, harness/API settings undisclosed | [stop → NEI] |
* Either publicly or privately to reviewers.
This is the minimum standard required to be Verified; failing any of these items results in a Flawed verdict.
| Scoring | |
|---|---|
| Examples |
|
| Default threshold for Flawed | ≥20% of inspected sample† contains errors or there is an issue that corrupts grading at scale |
| Status | PassFlag [stop → Flawed]Not reviewed |
| Notes | — |
| Benchmark Consistency | |
|---|---|
| Examples |
|
| Default threshold for Flawed | Leaderboard has incomparable results from different versions |
| Status | PassFlag [stop → Flawed]Not reviewed |
| Notes | — |
| Elicitation | |
|---|---|
| Examples | Model elicitation is extremely constraining and is not the focus of the benchmark:
|
| Default threshold for Flawed | Substantially reduced performance compared to reasonable alternatives for the tasks |
| Status | PassFlag [stop → Flawed]Not reviewed |
| Notes | — |
| Bias in Evaluation Setup | |
|---|---|
| Examples |
|
| Default threshold for Flawed | Material model-specific advantage found |
| Status | PassFlag [stop → Flawed]Not reviewed |
| Notes | — |
† For benchmarks with >50 available tasks, we will sample a random set of 50 (stratified by category, when present). If the observed error rate is 15–25%, we will expand the sample to 100. For benchmarks with ≤50 available tasks, all available tasks will be assessed.
This is the standard we would like all benchmarks to meet, but it is not necessarily disqualifying to omit or fail these items.
| Question | Status | Notes |
|---|---|---|
| Elicitation and resource adequacy: Are the resources given to models (reasoning token/turn budget, tool access, etc.) sufficient for them to perform near their ceiling? | SufficientConstrainingUnreasonably constrainingUnknownNot reviewed | — |
| Scaffold fairness: What scaffold does the leaderboard report? | Shared common scaffoldMix of model-specific and common scaffoldsModel-specific scaffoldsNot reviewed | — |
| Is there evidence/risk of contamination? | As of review date: —% of tasks public, —% of solutions publicNot reviewed | — |
| Has human completability been assessed? | All tasksRepresentative set of tasksPoor implementation (Unrepresentative set of tasks, unreasonable set of participants)Not establishedNot reviewed | — |
| Score range (if possible to estimate) | — floor, — ceilingNot reviewed | — |
| Statistical adequacy: how many runs/model (≥ 5 recommended for error bars) | — runs/modelUnknownNot reviewed | — |
| Construct Validity | Measures stated capabilitiesPartially measures stated capabilitiesDoes not measure stated capabilitiesNot reviewed | — |
For benchmarks with existing credible reviews, we will cite the external review rather than replicate the findings.