This work was supported by Schmidt Sciences.
Interested in purchasing access to solution verifiers? See below.
A Genus 2 Curve over the Rationals with a Rational Torsion Point of Prime Order at Least 31
Short Superpermutations over 8, 9, and 10
Large ℓ-Rank in Class Groups of Imaginary Quadratic Fields
Existence of EFX Allocations
Problems in FrontierMath: Open Problems are designed so that, even though no solution is known today, potential solutions can be checked for accuracy by a bespoke computer program, which we call a verifier.
Access to the verifiers is available for purchase by any party. Proceeds help fund the expansion of the benchmark. Our main cost is compensation to mathematicians, as the problems and verifiers are labor-intensive to formulate and implement.
At present, OpenAI is the only entity to have purchased access to the verifiers. OpenAI funded the creation of the original FrontierMath: Tiers 1-4, but Open Problems is developed independently and owned solely by Epoch.
Contact us with inquiries about purchasing access to the verifiers.
2026-08-11: We have marked the inverse Galois problem for the Mathieu Group \(M_{23}\) as solved by humans, as this problem has been solved by Huang et al. We hold a high bar for considering a problem “solved by AI”, in particular requiring that the core ideas of the solution be unambiguously contributed by AI. To judge this, we rely heavily on author attribution. In this case, one author is quoted as saying, “The boundaries between the human- and AI-contributed reasoning are not clear. Probably, it would be practically impossible to draw a sharp line.” Another author’s reflections state that, “Some of the most important decisions were clearly mathematical judgments made by the human collaborators.” We read this as saying it is not unambiguously clear that AI contributed the core ideas, and thus we do not consider this problem to be “solved by AI”.
2026-07-31: We have expanded the benchmark to 50 problems. We have also removed two problems from the benchmark, one about finding a surface with a high number of singularities and the other about finding an algorithm to decide whether a knot has unknotting number equal to 1. This was because we determined that the verifiers for these problems would not detect correct solutions with high enough fidelity.
We also removed the Ramsey-style hypergraph problem from the benchmark. Here we determined that, in hindsight, the problem did not meet our minimum bar for notability.
2026-06-27: We have stopped displaying failed AI attempts on the website. Problem pages will continue to be updated to reflect notable partial progress and solutions. We also modified the prompt for the problem about finding a surface with a high number of singularities. The verifier for this problem needs to be modified before it can handle what was described as Method C.
2026-03-05: We removed one problem from the benchmark, as we have determined that any solution would not meet our bar of being a publishable result in its own right. The problem page remains up: see it for more info on an AI-generated solution and subsequent human elaboration.
2026-02-24: We added two problems to the benchmark: finding a Hadamard matrix of order 668 and proving that certain “small” Diophantine equations have infinitely many solutions..
We have marked the inverse Galois problem for the Mathieu Group \(M_{23}\) as solved by humans, as this problem has been solved by Huang et al. We hold a high bar for considering a problem “solved by AI”, in particular requiring that the core ideas of the solution be unambiguously contributed by AI. To judge this, we rely heavily on author attribution. In this case, one author is quoted as saying, “The boundaries between the human- and AI-contributed reasoning are not clear. Probably, it would be practically impossible to draw a sharp line.” Another author’s reflections state that, “Some of the most important decisions were clearly mathematical judgments made by the human collaborators.” We read this as saying it is not unambiguously clear that AI contributed the core ideas, and thus we do not consider this problem to be “solved by AI”.
We have expanded the benchmark to 50 problems. We have also removed two problems from the benchmark, one about finding a surface with a high number of singularities and the other about finding an algorithm to decide whether a knot has unknotting number equal to 1. This was because we determined that the verifiers for these problems would not detect correct solutions with high enough fidelity.
We also removed the Ramsey-style hypergraph problem from the benchmark. Here we determined that, in hindsight, the problem did not meet our minimum bar for notability.
We have stopped displaying failed AI attempts on the website. Problem pages will continue to be updated to reflect notable partial progress and solutions. We also modified the prompt for the problem about finding a surface with a high number of singularities. The verifier for this problem needs to be modified before it can handle what was described as Method C.
We removed one problem from the benchmark, as we have determined that any solution would not meet our bar of being a publishable result in its own right. The problem page remains up: see it for more info on an AI-generated solution and subsequent human elaboration.
We added two problems to the benchmark: finding a Hadamard matrix of order 668 and proving that certain “small” Diophantine equations have infinitely many solutions..
Have a question? Noticed something wrong? Let us know.
A collection of unsolved mathematical problems designed to test AI systems' ability to advance human mathematical knowledge.