AI in math

Few domains test AI reasoning as clearly as mathematics, where answers can be verified automatically and the hardest problems extend to the frontier of human knowledge. Epoch tracks how AI is performing on mathematical tasks over time, including through FrontierMath, our own benchmark of expert-level problems designed to test the limits of what today's best systems can do.

Filter

Type
In August, 25% of math preprints acknowledged AI use, up from 4% in April
Data Insight
Sep. 18, 2026
In August, 25% of math preprints acknowledged AI use, up from 4% in April

Acknowledgments of AI use in arXiv math preprints rose from 4% in April 2026 to 25% in August, with 6% crediting AI with a substantial research contribution.

By Tara Abrishami

GPT-6 Astra leads on math benchmarks, but not on software engineering
Data Insight
Sep. 16, 2026
GPT-6 Astra leads on math benchmarks, but not on software engineering

OpenAI's GPT-6 Astra tops the Epoch Capabilities Index (ECI) with a score of 166, ahead of Claude Fable 5.1 at 164 and GPT-5.6 Sol at 162. Its Math-ECI of 170 sets a new record, but on software engineering benchmarks its SWE-ECI of 164 still lags behind Fable 5.1's 167.

By Alexander Barry and Jaeho Lee

Announcing FrontierMath Erdős
Update
Sep. 1, 2026
Announcing FrontierMath Erdős

A benchmark of 68 significant Erdős problems, open as of August 2026, curated by Thomas Bloom and formalized in Lean. AI systems must prove or disprove them within a fixed budget.

By Tom Adamczewski and Greg Burnham

Claude overperforms at software engineering and underperforms at math
Data Insight
May 15, 2026
Claude overperforms at software engineering and underperforms at math

Relative to their general Epoch Capabilities Index (ECI) values, Anthropic’s Claude models overperform on software engineering benchmarks (aggregated by the SWE-ECI) and underperform on math (Math-ECI). The SWE overperformance has been consistent across most generations, and remains in recent models. The math gap may be narrowing — Opus 4.6 and 4.7 both have Math-ECIs within 1 point of their general ECI, compared to larger gaps for earlier models.

By Alexander Barry

RIP Classic Reasoning Benchmarks. What's Next?
Newsletter
May 5, 2026
RIP Classic Reasoning Benchmarks. What's Next?

Give up at least one of: text only, short time horizon, easy to grade, and expert human superiority.

By Greg Burnham

Are AI benchmarks doomed?
Podcast
May 1, 2026
Are AI benchmarks doomed?

In this episode, Greg Burnham and Tom Adamczewski join Anson Ho to push back on benchmark pessimism and dig into what the next generation of AI benchmarks could look like.

By Greg Burnham, Tom Adamczewski, and Anson Ho

AI math capabilities could be jagged for a long time – Daniel Litt
Podcast
Jan. 29, 2026
AI math capabilities could be jagged for a long time – Daniel Litt

In this episode, Daniel Litt chats with the hosts about AI’s limits in mathematics, accelerating math research, and how to measure progress on open problems.

By Daniel Litt, Greg Burnham, and Anson Ho

Benchmarking AI on unsolved math problems
Update
Jan. 27, 2026
Benchmarking AI on unsolved math problems

Benchmarking AI on a collection of unsolved mathematics problems that have resisted serious attempts by professional mathematicians.

By The Epoch AI Team

Less than 70% of FrontierMath is within reach for today’s models
Newsletter
Oct. 17, 2025
Less than 70% of FrontierMath is within reach for today’s models

57% of problems have been solved at least once.

By Greg Burnham

Evaluating Gemini 2.5 Deep Think's math capabilities
Report
Oct. 9, 2025
Evaluating Gemini 2.5 Deep Think's math capabilities

Improved use of knowledge and precision, helpful for research, more conceptual in geometry, but limited creativity and citation issues.

By Greg Burnham