Our database of benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks administered internally by Epoch AI as well as data collected from external sources. Explore trends in AI capabilities across time, by benchmark, or by model.
The ECI combines scores from many different AI benchmarks into a single “general capability” scale.
We track information on pricing, architecture, and training, and show how models compare across benchmarks.
Featuring both Epoch-created and external evaluations, covering mathematics, coding, and more.
Claude Opus 5.5 took the top spot on the Epoch Capabilities Index with a score of 167, narrowly ahead of GPT-6 Astra. Claude Sonnet 5.5 roughly matched Claude Fable 5.1 (165).
GPT-6 Astra set new records on the ECI, as well as our math, continual learning, and game-puzzle benchmarks. OpenAI gave us pre-release access to test the model.
We've launched FrontierMath Erdős: 68 unsolved Erdős problems, with AI systems tasked with writing solutions in Lean. No prior model solved any of them; GPT-6 Astra solved 2 of 68, scoring 3%.
Need deeper insights? Our team offers custom research and advisory services.
Book a consultation