We fit the model to benchmark scores covering 59 benchmarks spanning 2023–present, drawn from the Epoch Benchmarking Hub, and currently report ECI scores for 264 models. The lists below are generated from the same configuration that drives the ECI pipeline, so they always reflect the benchmarks in the current fit.
Chess Puzzles, Earthborne Rangers (EBR-bench), FrontierMath Tier 4 (v1, old), FrontierMath Tier 4 (v2), FrontierMath Tiers 1-3 (v1, old), FrontierMath Tiers 1-3 (v2), GPQA Diamond, MATH Level 5, MirrorCode, Mystery Game Puzzles, OTIS Mock AIME 2024-2025, SimpleQA Verified, SWE-bench Verified
Aider Polyglot, APEX-Agents, ARC-AGI-1, ARC-AGI-2, BALROG, CadEval, CL-bench, CL-bench Life, DeepResearchBench, DeepSWE, DTBench, ExploitBench, Fiction.liveBench, FrontierCode, GDPval, GeoBench, GSO, Humanity's Last Exam, Lech Mazur Writing, LMCA, METR Time Horizons, OS World, OSWorld 2.0, PostTrainBench, ProofBench, Remote Labor Index, SimpleBench, Surface Evolver Bench, Terminal-Bench 2.0, The Agent Company, VPCT, WeirdML (v2)
Adversarial NLI, ARC AI2, BBH, Cybench, GSM8K, HellaSwag, LAMBADA, MMLU, OpenBookQA, PIQA, ScienceQA, SuperGLUE, TriviaQA, WinoGrande
To be used in our methodology, benchmarks need to be scored on a 0-1 scale (or equivalently, 0% to 100%). For benchmarks where random guessing would score above 0 (e.g. those with multiple choice responses), we rescale so that random guessing performance is scaled to zero. For benchmarks with a known error rate, we also rescale scores so that the maximum possible score is scaled to one. For example, FrontierMath v1 was found to have an error rate of 43% across Tiers 1-3, and 40% for Tier 4, so its scores are rescaled accordingly. Where a model has been evaluated on both v1 and v2 of a FrontierMath benchmark, only the v2 score enters the fit.
In order to capture the upper end of each model’s capabilities, we aggregate across model evaluation settings (e.g. thinking effort and inference provider), taking the highest score for each benchmark. We only aggregate over models released on the same day with the same name (e.g. we do not aggregate across versions of GPT-4o released on different days). We drop any models with fewer than 4 benchmark scores from our model fit, to avoid low-certainty estimates. We also exclude models released before 2023 due to data sparsity; we hope to increase coverage in the future.