Data Insight
Sep. 16, 2026

GPT-6 Astra leads on math benchmarks, but not on software engineering

OpenAI’s GPT-6 Astra tops the Epoch Capabilities Index (ECI), our composite measure of model capabilities, with a score of 166, ahead of Anthropic’s Claude Fable 5.1 at 164 and OpenAI’s own GPT-5.6 Sol at 162. But this hides some complexity: while its Math-ECI of 170 sets a new record, on software engineering benchmarks Astra’s SWE-ECI of 164 still lags behind Fable 5.1’s 167.

Based on pre-release evals we gave Astra an ECI of 169 when it launched, but this fell as more software engineering benchmark results became available. Math-ECI and SWE-ECI are both domain-specific ECIs, refit on only one domain’s benchmarks, measuring a model’s strength relative to the general ECI. Explore these and custom subsets in our Domain-specific ECI Explorer.

Epoch's work is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons BY license.

Learn more about this graph

We compare four recent models (GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol and Kimi K3) using the Epoch Capabilities Index (ECI), as well as domain-specific ECI variants meant to measure math and software capabilities. Domain-specific ECIs retain the benchmark difficulty parameters from the general ECI fit and re-estimate each model’s capability from only a subset of those benchmarks from a specific domain, making the resulting scores comparable to general ECI scores. See our domain-specific ECI documentation for details on methodology.

Download this data

Explore this data