FAQ
What do you hope to achieve by doing this?
We hope to inform the public about the quality of notable benchmarks, and to raise the standard of benchmarks. To date, no systematic quality check is done, and benchmark audits have often revealed extensive errors. Existing reviews provide valuable information, but are ad-hoc and do not follow a common format. We see an opportunity for Epoch to become a reliable source of benchmark reviews and a resource for benchmark creators across industry and academia.
Our main goals are to:
- Inform the public about benchmark quality
- Incentivize benchmark developers to create high-quality benchmarks
- Increase the information available about how to create good benchmarks
Why does this matter?
Benchmarks are the main way to publicly assess how capable AI systems are. A deeper understanding of what we are measuring and what it represents informs our understanding of AI capability. Errors in benchmarks can bias leaderboard results and lead to misinterpreted results. This is important for deployment outcomes and policy decisions.
A broad set of benchmark reviews helps us evaluate the ecosystem as a whole and understand what we are measuring well and what we are missing. We also hope to increase information about common types of errors in order to raise the bar for all benchmarks.
Increasingly, environment-level issues lead to reward hacking behaviors that affect numerous benchmarks, especially agentic benchmarks1. Significant variance is also a result of scaffold specifications. In our prior analysis, “Why benchmarking is hard,” we estimated that switching the scaffold makes up to an 11% difference for GPT-5 and up to a 15% difference for Kimi K2 Thinking. Finally, when specifications such as resource limits and harness settings are poorly chosen or unspecified, evaluators may run the same benchmark under different conditions, making meaningful comparisons difficult.
What does Verified mean?
Verified benchmarks can broadly be interpreted as described. They may still contain issues that do not substantially affect the results, and these will be noted in the review. The minimum standards to achieve a Verified label are listed in our rubric.
What does Flawed mean?
Flawed benchmarks have at least one major flaw that affects their interpretability. Representative errors and the threshold for a benchmark to be considered flawed are listed in the rubric. Failing to meet the standard in any of the categories results in a Flawed verdict.
My benchmark was labeled Flawed. If I fix the errors, will my benchmark be Verified?
Our reviews reference specific snapshots of benchmarks, so the version we reviewed is static and its label will remain unless there is an issue with the content of the review itself. If you publish a new version that addresses the listed errors, we will consider it for re-review, but it may still not meet the standards to be labeled Verified. Once we identify sufficient errors for a flawed review, we stop reviewing and publish our findings. We do not do a comprehensive review of the entire benchmark. There may thus be additional errors we did not identify in our first review. We hope to review new versions of benchmarks, but are limited in capacity.
Why isn’t my benchmark included?
We are constantly expanding the list of benchmarks we have reviewed. Our prioritization is described in the “Included benchmarks” section.
Why wasn’t I given the opportunity to fix my benchmark before publication?
We think it’s vital to get information to the public as quickly as possible, especially for highly cited benchmarks. We reach out to all developers at least one day prior to publication to give notice and will publish the developer’s response in full, if they desire.
I disagree with the verdict that my benchmark is flawed. What should I do?
Email us at reviews@epoch.ai.