Our database of benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks administered internally by Epoch AI as well as data collected from external sources. Explore trends in AI capabilities across time, by benchmark, or by model.
The ECI combines scores from many different AI benchmarks into a single “general capability” scale.
We track information on pricing, architecture, and training, and show how models compare across benchmarks.
Featuring both Epoch-created and external evaluations, covering mathematics, coding, and more.
GPT-6 Astra set new records on the ECI, as well as our math, continual learning, and game-puzzle benchmarks. OpenAI gave us pre-release access to test the model.
We've launched FrontierMath Erdős: 68 unsolved Erdős problems, with AI systems tasked with writing solutions in Lean. No prior model solved any of them; GPT-6 Astra solved 2 of 68, scoring 3%.
We ran a human baseline on Earthborne Rangers, the board game underlying EBR-bench. Top human players reached mastery after five playthroughs.
Need deeper insights? Our team offers custom research and advisory services.
Book a consultation