Benchmark Reviews Documentation

OverviewMethodologyIncluded benchmarksFAQ
Show sidebarIncluded benchmarks

Included benchmarks

For the initial release of Benchmark Reviews, we attempted to create a representative sample of benchmarks across error types and domains. Benchmarks that require review by experts with domain knowledge take additional time as we hire external subject matter experts.

We plan to prioritize reviews with the following characteristics going forward:

  • Impact: Evaluates a task we care about and/or safety-relevant capabilities
  • Reach: Widely cited, in recent system cards, or underrated
  • Diversity: Covers an under-covered capability in our current suite of benchmark reviews

Many benchmarks do not publicly release enough information for us to be able to sufficiently evaluate them against our rubric. These will be marked as Not enough information. For benchmarks where we have been given access to private information, we will disclose what we were able to evaluate. In order to avoid leakage, we will only include generic information in these reviews, such as scoring error types without task-level analysis.

We will not review Epoch-created benchmarks due to conflict of interest. This does not mean we think our benchmarks are without errors1. We may commission external review of our benchmarks in the future. We generally do not expect to be compensated for reviewing benchmarks, and we will clearly note situations where this is not the case.

Reviewed benchmarks

BenchmarkVerdictReview date
ExploitBench v0.1 VerifiedSep. 12, 2026
CritPt Not enough infoSep. 11, 2026
FrontierCode Not enough infoSep. 11, 2026
Berkeley Function Calling Leaderboard (BFCL) v4 FlawedSep. 10, 2026
HealthBench Professional FlawedSep. 10, 2026
SimpleQA Verified VerifiedSep. 10, 2026
PostTrainBench v1.1 VerifiedSep. 9, 2026
DeepSWE v1.1 FlawedSep. 7, 2026
Terminal-Bench 4.0.0 FlawedSep. 4, 2026
SWE-bench Verified FlawedSep. 3, 2026
SWE-Bench Pro FlawedSep. 1, 2026
Humanity's Last Exam FlawedAug. 11, 2026
Lech Mazur Writing FlawedAug. 10, 2026
TextQuests FlawedAug. 10, 2026
WeirdML v2 VerifiedAug. 10, 2026
Notes
  1. An audit of FrontierMath v1 found 42% of problems contained errors. https://x.com/EpochAIResearch/status/2065488154086568445 Return