Robust research-level mathematical reasoning by LLMs
Establish robust research-level mathematical reasoning capabilities in large language models by achieving significantly improved performance on unpublished, research-level benchmarks such as FrontierMath, thereby closing the current gap indicated by low accuracy on this benchmark.
References
Similarly, on the FrontierMath Benchmark, which consists of unpublished research-level problems, Gemini 3 Pro scores 18.75% on the research-level split, indicating that robust research-level reasoning remains an open problem.
The answer to Question \ref{crq} depends crucially on a subquestion about what AI tools will actually be able to do. It is convenient to formulate this pseudomathematically, not as a single conjecture, but as a family of conjectures indexed by a large number of free parameters. At some point in the near future, some AI tools will, at some expense, and with some level of human supervision, be able to accomplish some research-level mathematical tasks in some fields of mathematics, with some non-trivial success rate, and at some level of correctness and quality.