Third Batch Benchmark

The Third Batch Benchmark will test AI systems on ten naturally occurring research math problems which have been solved by human mathematicians but whose solutions have not been made public.

If you are interested in having your AI system tested in the benchmark, read the call for submissions and fill in this form by September 10, 2026.

Timeline

Problem selection

Mathematicians across fields submit unpublished problems with proofs of at most 12 pages. All submissions undergo a first round of refereeing, and 10 problems are selected for the benchmark.

Benchmark testing

The editorial board tests AI systems via API. Each system is an open source harness which calls publicly available models, and gets one shot per question with no additional interaction.

Benchmark grading

Roughly 30 expert referees review the AI solutions double blind, scoring correctness, exposition, and attribution separately. Results — including the problems, human solutions, referee reports, and editorial board scores — will be published on October 14, 2026.