Our database of benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks evaluated internally by Epoch AI as well as data collected from external sources. Explore trends in AI capabilities across time, by benchmark, or by model.
The ECI combines scores from many different AI benchmarks into a single “general capability” scale.
We track information on pricing, architecture, and training, and show how models compare across benchmarks.
Featuring both Epoch-created and external evaluations, covering mathematics, coding, and more.
We've updated the MirrorCode leaderboard. Claude Fable 5 leads with a score of 64%, followed by GPT-5.6 Sol at 20%.
We've launched an expansion of FrontierMath: Open Problems! The benchmark now contains 50 significant, unsolved problems from research mathematics. AI has solved three so far.
AI has found a presentation for the absolute Galois group of the field of 2-adic numbers — the second problem solved in FrontierMath: Open Problems.
Need deeper insights? Our team offers custom research and advisory services.
Book a consultation