Benchmarks

A registry of 80 AI benchmarks covering mathematics, software engineering, agentic workflows, games, and more.

Filter

Creator
Highest score
Administered by
Domain
80 results

An index aggregating many different benchmarks into a single, general capability scale.

266 models tested
Epoch administered

A collection of unsolved problems from research mathematics, curated to be especially interesting and difficult.

Mathematics
Epoch administered

A collection of problems posed or studied by mathematician Paul Erdős, open as of August 2026 and curated to be especially interesting and difficult.

Highest score: 3%
MathematicsAgent
5 models tested
Epoch administered
MirrorCode By Epoch

A long-horizon coding benchmark. Models reimplement entire programs end-to-end, without access to the original source code.

Highest score: 73%
Software engineeringAgent
8 models tested
Epoch administered

A test of AI systems' learning capabilities, measuring whether their scores improve across repeated playthroughs of the relatively obscure game Earthborne Rangers.

Highest score: 76%
GamesAgent
21 models tested
Epoch administered

100 positions from a puzzle-oriented variant of a well-known game, the details of which are deliberately left undisclosed.

Highest score: 84%
Games
127 models tested
Epoch administered
Chess Puzzles By Epoch

100 novel puzzles, generated programmatically with a chess engine. Each puzzle has a single best next move.

Highest score: 72%
Games
222 models tested
Epoch administered

1,000 factoid questions about politics, science and technology, art, sports, geography, music, and more.

Highest score: 76%
World knowledge
80 models tested
Epoch administered

Exceptionally difficult research-level math problems.

Highest score: 98%
MathematicsAgent
62 models tested
Epoch administered

Expert-written math problems covering advanced undergrad through early career research.

Highest score: 94%
MathematicsAgent
106 models tested
Epoch administered