The editorial board helped Epoch curate problems for the benchmark. Their overall comments on the project are below.
Royal Society University Research Fellow, University of Manchester
The problems on the FM:OP benchmark are all at least reasonably difficult, and of interest to at least some mathematicians. This sounds like a low bar, but is the best the one can realistically hope for, as mathematical tastes differ widely in the community (and the editorial board). No doubt every mathematician will see some problems on this list and think “why on earth did they include that?” — but as long as they also see others and think “naturally that should be included, it’s a deep and important problem”, then I think we have done a good job.
I expect that progress on this benchmark will be uneven; with the release of each new model, new classes of problem become tractable, and hence we will probably see batches of solutions, then long periods of inactivity. I also expect that we will see some solutions, not from an AI alone, but from human researchers working in collaboration with an AI — and, of course, hope that there will still be some progress made by humans alone!
That the problems have all been the subject of some serious study hopefully means they are all at least somewhat difficult; there is no upper bar on the difficulty of the problems. They have been chosen so that, with reasonable probability, they are outside of the range of current techniques, so to make meaningful progress on many of them, substantial new ideas are required.
This is not a guarantee, or a promise. Many times in the history of mathematics we have been wrong in our estimation of the difficulty of a problem, often after the discovery of a surprisingly elegant and simple proof. It is possible that the difficulty of the some of these benchmark problems may similarly have to be adjusted with hindsight (if not officially, then at least in my view) if the solution is simpler than expected.
It is possible (although this would be a little disappointing) that some (hopefully not many) of the problems will be solved via a combination of clever brute force and existing, perhaps obscure, ideas in the literature. No doubt even this would be of interest to those mathematicians working in the relevant area, and there is immense value in discovering new connections to existing work from other areas. (This has happened with the recent counterexample to the unit distance conjecture for example, which has led to the addition of deep results from class field theory to the toolkit of the combinatorialist.)
The hope, however, is that many of these problems are difficult enough that an AI will have to invent new techniques to make progress. This is the large question that looms over the intersection of AI with research mathematics: can an AI develop new techniques to solve a difficult open problem?
Thus far the progress made by AI has been made using existing ideas and techniques, often in creative and unexpected ways. The extreme specialisation of mathematics means that it has developed in a very spiky fashion, and there are many gaps in the convex hull of what we know, waiting to be filled in — AI has proved very capable at drawing connections to help with this task, which is already invaluable in research.
But can an AI genuinely push out the boundary of this convex hull, and break new ground? Collecting established, difficult, problems as in this benchmark will help us measure whether this is happening, and will help guide the efforts of AI in a direction that is interesting to mathematicians.
How much AI has succeeded on this benchmark will not be captured by a single number, or collection of numbers. The true measure of progress will be more qualitative, and I encourage those interested to not just view the scorecard, but to read and reflect on the methods used, and the commentary of experts on such. Are these methods new, or a synthesis of existing ideas? Do the methods lead to the solution of other problems? Do they lead to a fundamental change in how we understand an area, or are they a collection of ad hoc tricks useful only for a single purpose? As with human mathematical research, the true importance and consequences of a solution to any given problem may take years or decades to become evident.
If the benchmark is saturated, and all of the problems on it are solved, that would mean that AI has reached or gone beyond the frontier of research mathematics in a wide range of areas, for a certain type of question. It does not necessarily mean that it is superhuman in every aspect of research; there may be something about the kind of constructions asked for that make them particularly amenable to AI progress. Finding concrete constructions for well-defined questions is only a part (albeit an important part) of research mathematics — and, of course, once any problem is solved, there is a more difficult variant lurking on the horizon.
Finally, a note on formalised proofs. I’m sure many will be wondering “Why is this benchmark limited to problems with an easily verifiable solution? Why not have a more typical list of open problems, stated formally in Lean, and ask for formalised proofs?” Lists of this type are also an important way to measure AI progress and to set concrete goals, and I don’t want to suggest that this benchmark is superior to them, but rather plays a complementary role. Indeed, I hope that there will soon emerge some other benchmarks of formalised conjectures, with some kind of organisation and curating of problems as in this benchmark, rather than the ad hoc approach taken so far.
This benchmark collects problems with the unusual restriction that they must have an easily verifiable answer. This means that we can decouple measuring the ability of AI to produce interesting mathematics (even if this is done in some inscrutable fashion that takes humans a long time to puzzle out) from the ability of AI to correctly formalise its arguments. In many areas formalising an argument is much more difficult than coming up with an intuitive path that seems to work. I am interested to see how well AI can do when asked to explore freely, without being constrained by the need for formalised proofs.
Assistant Professor of Mathematics, University of Toronto
I’m quite excited about FM:OP’s potential to measure the capabilities of AI to do interesting mathematics. We’ve recently seen a flood of AI-generated mathematical work, much of which has not yet been checked for correctness (let alone checked for interesting new ideas, with more applications). While it’s clear that AI capabilities are rapidly increasing, and they are starting to do genuinely interesting new mathematics, measuring this increasing capability is also increasingly labor-intensive and unreliable. It is costly to generate interesting questions, and exceedingly labor-intensive to verify their solutions.
FM:OP problems are taken from the literature, evaluated for interest, and, crucially, have verifiable solutions. This is not to say that generating these problems and writing their verifiers is cheap or easy, but it goes some way to addressing the issues above. What’s more: there is pretty interesting mathematics involved in massaging an open question into a programmatically verifiable format. I hope my colleagues take the opportunity to think about which of their favorite conjectures admit such forms.
It’s a messy project, fitting for a messy new age of mathematics. The format has problems: by the nature of the verifiability requirement, the benchmark will miss vast swathes of mathematical activity. Some of the problems may be unanswerable; we’ll never know when the benchmark is saturated. My expectation is that there will be a spate of “easy” problems (and perhaps some not-so-easy ones) resolved over the next couple of months and years, and then a lull.
The problems span a broad range of interestingness. It’s possible that some have solutions that are, in some sense, implicit in existing literature. Some I frankly do not care for (though I am willing to accede to the opinions of specialists). For others, I think solutions would likely introduce important new ideas, and almost certainly be interesting. A number of the problems have the form “break the following existing record.” These problems (and many others) can only really be evaluated in retrospect: did the model solution do something new? Or was it just scaling up existing ideas, or compute? Some of the problems may turn out to have verifiers that are measuring something weaker than what the problem is asking for, or that have some vulnerability a model can find. Human review will still be essential.
What will we learn? Something about AI capabilities — my view is that this is much more easily understood in retrospect. Chiefly, I hope, some interesting math.
Professor of Mathematics, University of California, Davis
FrontierMath: Open Problems is an ambitious effort to evaluate AI models’ performance on open research problems that are of interest to working mathematicians. In a departure from earlier benchmarks, the open status of the benchmark problems stands as a strong, honest signal of their difficulty level. Any success by an AI on these problems will represent a credible demonstration of reasoning abilities comparable to or potentially exceeding those of professional research mathematicians. Equally (or more) importantly, it will be a tangible contribution to the research enterprise; this makes it an “organic,” not “synthetic”, benchmark.
While not all the FM:OP problems are equally difficult or important, some are believed to be extremely difficult: success on one or more of the benchmark problems having the highest difficulty classification may have profound implications for mathematics research, suggesting that AI has legitimately earned a seat at the table as a collaborator capable of advancing humanity’s mathematical knowledge in new and interesting ways.
It’s critical to stress that even this most dramatic type of (hypothetical) progress would not mean that AI has surpassed humans in overall mathematical research capabilities. The benchmark does not purport to measure all of the reasoning skills mathematicians rely on, and its design limits it to problems with specific structural characteristics, and to goals that favor the finding of counterexamples over establishing that all objects satisfy a certain property. (For example, disproving the Riemann Hypothesis would be a type of problem that satisfies the benchmark’s verifiability criteria, while proving the Riemann Hypothesis would not qualify.)
Looking forward on the trajectory of AI mathematical reasoning, I expect that we will all continue to be surprised and (mostly) delighted by the new emerging AI capabilities in mathematical reasoning. I predict that we will see at least one genuinely dramatic new mathematical result proved by AI in the next year. On the other hand, for the next year at least, I expect that AI will continue to be quite limited in the types of mathematical breakthroughs it can come up with: any breakthrough proofs will be fairly short, and will skew heavily in the direction of the discovery of counterexamples or “rare” mathematical objects.
The problems in the benchmark are all interesting — I was rather inspired by the diversity of research areas and topics covered, which taught me a lot about what fascinating things people are working on these days. Even so, FM:OP covers only a small fraction of the true landscape of mathematical research. Mathematicians should not wait until AI companies get around to testing models’ capabilities on their favorite research areas, but would be wise to engage with this exciting frontier through active exploration of their own.