Included benchmarks
For the initial release of Benchmark Reviews, we attempted to create a representative sample of benchmarks across error types and domains. Benchmarks that require review by experts with domain knowledge take additional time as we hire external subject matter experts.
We plan to prioritize reviews with the following characteristics going forward:
- Impact: Evaluates a task we care about and/or safety-relevant capabilities
- Reach: Widely cited, in recent system cards, or underrated
- Diversity: Covers an under-covered capability in our current suite of benchmark reviews
Many benchmarks do not publicly release enough information for us to be able to sufficiently evaluate them against our rubric. These will be marked as Not enough information. For benchmarks where we have been given access to private information, we will disclose what we were able to evaluate. In order to avoid leakage, we will only include generic information in these reviews, such as scoring error types without task-level analysis.
We will not review Epoch-created benchmarks due to conflict of interest. This does not mean we think our benchmarks are without errors1. We may commission external review of our benchmarks in the future. We generally do not expect to be compensated for reviewing benchmarks, and we will clearly note situations where this is not the case.
Reviewed benchmarks
| Benchmark | Verdict | Review date |
|---|---|---|
| ExploitBench v0.1 | Verified | Sep. 12, 2026 |
| CritPt | Not enough info | Sep. 11, 2026 |
| FrontierCode | Not enough info | Sep. 11, 2026 |
| Berkeley Function Calling Leaderboard (BFCL) v4 | Flawed | Sep. 10, 2026 |
| HealthBench Professional | Flawed | Sep. 10, 2026 |
| SimpleQA Verified | Verified | Sep. 10, 2026 |
| PostTrainBench v1.1 | Verified | Sep. 9, 2026 |
| DeepSWE v1.1 | Flawed | Sep. 7, 2026 |
| Terminal-Bench 4.0.0 | Flawed | Sep. 4, 2026 |
| SWE-bench Verified | Flawed | Sep. 3, 2026 |
| SWE-Bench Pro | Flawed | Sep. 1, 2026 |
| Humanity's Last Exam | Flawed | Aug. 11, 2026 |
| Lech Mazur Writing | Flawed | Aug. 10, 2026 |
| TextQuests | Flawed | Aug. 10, 2026 |
| WeirdML v2 | Verified | Aug. 10, 2026 |
-
An audit of FrontierMath v1 found 42% of problems contained errors. https://x.com/EpochAIResearch/status/2065488154086568445