Mathematics in the age of AI
Abstract: An essay, based on a public lecture delivered at the 2026 International Congress of Mathematicians, on how the mathematical community might respond to the arrival of artificial intelligence tools that are capable of performing research-level mathematical tasks. Rather than debating the capabilities of such tools, we condition on the hypothesis that these capabilities will arrive, and examine instead a question that is orthogonal to it: what the goals and values of mathematical research actually are. The problem-solving component of mathematics is used as a case study.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is the paper about?
This paper asks how mathematicians should respond if artificial intelligence becomes able to solve difficult, research-level mathematics problems.
The author does not mainly debate whether AI will become this powerful. Instead, he asks a different question:
If AI can solve many hard problems, what should mathematics value, and how should mathematicians work?
The paper argues that mathematics is not only about producing correct answers. It is also about understanding ideas, explaining them, checking them, teaching them, and building shared knowledge.
2. Main questions and objectives
The paper focuses on two connected questions.
Will AI be able to do advanced mathematics?
The author assumes, for the purpose of discussion, that AI will soon be able to solve a significant number of difficult mathematical problems. This is called the Working Hypothesis.
This does not mean the author is certain that it will happen. He is saying: Let us imagine that it does happen and think carefully about the consequences.
What are the real goals of mathematics?
Mathematics has many goals, including:
- solving difficult problems;
- creating new theories and methods;
- understanding the world;
- teaching future mathematicians;
- building a mathematical community;
- creating beautiful and lasting ideas;
- adding reliable knowledge to a shared collection of human understanding.
The paper argues that mathematicians have often acted as if solving problems quickly is the main goal. But this may only have worked because the other goals usually improved at the same time.
For example, solving one hard problem might also create a useful theory, train students, inspire other researchers, and lead to new applications. AI could break this connection if it produces answers much faster than people can understand or use them.
3. How does the paper approach the topic?
This is an essay, not a traditional experiment or data-based scientific study. It uses historical examples, logical analysis, and a detailed example about mathematical problem solving.
A historical comparison
The author compares the arrival of powerful AI with the crisis in the foundations of mathematics in the early twentieth century.
At that time, mathematicians discovered serious problems in the basic rules they were using. For example, Russell’s paradox showed that some simple ideas about collections, or “sets,” could lead to contradictions. Later, Gödel’s incompleteness theorems showed limits on what formal mathematical systems can prove.
These problems were difficult, but they eventually led to clearer and more reliable foundations for mathematics.
The author thinks AI may cause a similar period of change—not by threatening mathematical truth, but by forcing mathematicians to examine their unwritten rules and values.
The problem-solving “pipeline”
The paper studies problem solving as a series of stages. The author begins with a simple goal:
- Solve as many unsolved problems as possible.
He then shows why this is not enough. The process must include several additional steps:
- Generate a proof — find a possible solution.
- Verify the proof — check that every step is correct.
- Explain the proof — write it clearly enough for people to understand.
- Publish and review it — have experts examine and discuss it.
- Accept it into the community — other mathematicians decide that it is useful and trustworthy.
- Add it to the theory of the subject — connect it with other results and include it in textbooks or standard references.
This is like discovering a new machine. It is not enough for the machine to work once. People also need to understand how it works, check that it is safe, explain it to others, and decide how it fits into everyday life.
Technical terms in simple language
- A proof is a careful explanation showing why a mathematical statement must be true.
- Proof verification means checking every part of the proof.
- A proof assistant is a computer program that checks mathematical reasoning, much like a very strict grammar checker checks every sentence.
- Formalization means rewriting mathematics in a precise computer language so that a program can inspect it.
- Peer review is when other experts examine a piece of research before it is officially accepted.
- Canonicalization means turning a new result into a polished, standard part of mathematical knowledge—for example, presenting it in the best form in a textbook.
The paper also discusses Goodhart’s law:
When a measurement becomes the main target, it may stop being a good measurement.
For instance, if a school only rewards students for getting high scores, some students may focus on memorizing tricks rather than truly understanding the subject. Similarly, if mathematics rewards only the number of problems solved, people or AI systems may produce many shallow, confusing, or unhelpful results.
4. Main findings and why they matter
The paper’s main conclusions are arguments rather than experimental results.
Correctness is not enough
An AI might produce a proof that a computer confirms is correct, but that no human understands. The author argues that such a proof is not fully successful.
A proof should help mathematicians:
- learn a new idea;
- see why the result matters;
- use the method elsewhere;
- explain it to students;
- connect it with earlier mathematics.
A computer-checked answer can be like a locked box containing a treasure. The treasure may be real, but it is much less useful if nobody knows how to open or understand the box.
Clear writing is not enough either
An AI may produce a proof with perfect spelling and formatting, yet still fail to explain the important ideas. It might spend many pages describing easy steps while hiding the difficult and creative part.
The author also makes an interesting point about human writing. Human mathematical papers sometimes contain awkward parts, changes in notation, or especially careful explanations. These can show readers where the author struggled and where the important ideas are located. A perfectly smooth AI-written proof might remove these useful clues.
Mathematical progress depends on people
A mathematical result becomes truly valuable when other mathematicians understand it, trust it, teach it, and use it. This requires human activities such as:
- explaining ideas;
- checking sources;
- giving credit;
- reviewing papers;
- discussing results;
- teaching students;
- deciding which results belong in the standard theory.
AI may speed up the production of proofs, but it cannot automatically replace this larger social process.
Mathematics may face “proof abundance”
For a long time, difficult proofs were scarce. There were fewer results than mathematicians could study. If AI begins producing proofs very quickly, the situation could reverse.
There may be:
- more proofs than experts can check;
- more checked proofs than people can explain;
- more readable papers than journals can review;
- more published results than the community can absorb into textbooks and theories.
This is called a shift from proof scarcity to proof abundance.
The existing systems of mathematics—journals, referees, prizes, hiring practices, and academic recognition—were designed for a world where major results were relatively rare. They may not work well when large numbers of results can be created automatically.
5. Suggested responses and possible impact
The author does not offer one complete solution, but he recommends changing what mathematicians reward and value.
Be open about AI use
Researchers should clearly state when they used AI, proof assistants, or other automated tools. This allows readers and reviewers to understand how the work was produced and who is responsible for it.
Keep humans responsible
AI systems should not receive authorship or credit as if they were human mathematicians. Humans must remain responsible for the correctness, meaning, and proper attribution of a result.
Give proper credit
AI tools may fail to identify where ideas came from. Mathematicians therefore need to work carefully to find and credit earlier researchers.
Strengthen reviewing and explanation
The paper suggests that mathematical culture should give more credit to activities that help the community understand results:
- refereeing;
- writing clear explanations;
- creating surveys and textbooks;
- organizing mathematical databases;
- formalizing proofs;
- explaining the history and ideas behind discoveries.
The author’s practical rule is simple: if the authors cannot give a clear, expert-level explanation of their own result, it probably should not yet be published—even if a computer has checked the proof.
Protect education and human development
In teaching, AI must be used carefully. Producing a correct answer is not the same as learning mathematics. Students need to struggle with problems, make mistakes, develop intuition, and learn how to think.
Conclusion
The paper’s central message is that mathematics should not become a race to produce the greatest number of automatically generated proofs.
If AI becomes powerful at solving mathematical problems, it could be extremely useful. It might help discover new results, check long arguments, and explore ideas more quickly. However, the most important parts of mathematics—understanding, explaining, teaching, giving credit, and building a shared body of knowledge—will still require human judgment and cooperation.
The paper therefore encourages mathematicians to make their values more explicit. A successful mathematical result should not merely be correct. It should also be understood, communicated, trusted, shared, and connected to the wider theory of mathematics.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
The paper identifies important risks but leaves the following issues unresolved:
- The “Working Hypothesis” is not operationalized. Terms such as “reasonably soon,” “reasonable fraction,” “research-level,” “quality,” “success,” and “supervision” are not defined in measurable terms, making the conditional analysis difficult to test or apply.
- AI capabilities are not evaluated systematically. The discussion relies heavily on the First Proof project, whose sample contains only ten problems and may not represent different mathematical fields, difficulty levels, problem types, or real research workflows.
- The evidence base is vulnerable to selection bias. The paper notes reporting bias in AI demonstrations but does not provide a systematic review, benchmark protocol, failure taxonomy, or quantitative estimate of false positives and false negatives.
- The costs and resources required for AI-assisted mathematics remain unclear. The paper does not compare compute costs, energy use, human labor, supervision time, access to proprietary systems, or reproducibility across tools and institutions.
- The boundary between AI contribution and human contribution is unspecified. The paper does not establish how to attribute credit when humans choose problems, design prompts, develop intermediate ideas, correct errors, formalize arguments, or select among AI-generated approaches.
- The proposed human authorship principle lacks enforceable criteria. It remains unclear what level of understanding, explanation, responsibility, or intellectual contribution is sufficient for authorship and publication.
- “Understanding” is treated as a central requirement but is not defined or measured. The paper does not distinguish conceptual understanding, ability to reproduce a proof, ability to explain it pedagogically, or ability to generalize and adapt it.
- The claim that human-readable proofs are epistemically superior is not empirically tested. The argument about “natural friction,” tacit knowledge, and AI-polished exposition would benefit from studies measuring whether different exposition styles actually improve learning, error detection, or future research productivity.
- The proposed rule requiring an expert-level author talk is not evaluated. Such a requirement could exclude valid collaborative, highly technical, or interdisciplinary work, but the paper does not analyze its reliability, fairness, or practical enforceability.
- The five-stage problem-solving pipeline is conceptual rather than validated. The paper does not establish whether proof generation, verification, exposition, publication, acceptance, and canonicalization are genuinely separable stages or how outcomes at one stage affect the others.
- The scale of the predicted “proof abundance” is unknown. No model estimates how many AI-generated proofs will be produced, how many will be correct, or whether verification, exposition, peer review, and canonicalization will actually become bottlenecks.
- The effects of AI on peer review are not quantified. The paper proposes automated triage but does not assess its accuracy, susceptibility to gaming, effects on reviewer workload, or risk of filtering out unconventional but valuable mathematics.
- Formal verification is treated as more decisive than its practical limitations may allow. The paper does not examine the difficulty of formalizing advanced mathematics, gaps in existing libraries, dependence on trusted kernels and implementations, or the possibility that formal proofs can encode inappropriate definitions or irrelevant results.
- The relationship between formal correctness and mathematical significance remains unresolved. A proof assistant can establish validity but cannot by itself determine novelty, depth, explanatory value, generality, or importance.
- The paper does not analyze AI errors in attribution and literature integration in sufficient detail. It identifies attribution as a problem but does not propose methods for detecting fabricated citations, tracing conceptual provenance, identifying training-data contamination, or assigning credit for rediscovered ideas.
- Novelty is not clearly defined in an AI-assisted research environment. The paper does not address whether a result is novel when an AI system independently reproduces known mathematics, combines existing results, or generates a proof that was not previously documented.
- The impact on mathematical priority conventions is unexplored. Faster generation of candidate proofs may intensify disputes over discovery, prompting, unpublished human insights, and the ownership of AI-assisted results.
- The consequences for early-career mathematicians are underdeveloped. The paper does not determine how AI will affect training, apprenticeship, skill acquisition, employment, evaluation, or opportunities to develop mathematical intuition through struggle.
- Potential inequalities in access are not analyzed. Differences in compute, subscriptions, institutional support, data access, and technical expertise could widen disparities across countries, institutions, fields, and individual researchers.
- The paper does not examine conflicts between openness and commercial control. Proprietary models may limit reproducibility, auditability, disclosure, and long-term access to the systems used to produce mathematical results.
- The effects on mathematical diversity and research direction are left open. Optimizing benchmarkable goals may concentrate effort on tractable or well-represented areas while neglecting emerging fields, foundational questions, negative results, or culturally diverse mathematical traditions.
- Theory building receives only a brief acknowledgement. The paper does not analyze how AI may affect abstraction, conjecture formation, choice of definitions, identification of fruitful frameworks, or the development of theories that are not reducible to solving existing problems.
- Teaching, mentoring, hiring, grants, outreach, and refereeing are identified as important domains but not analyzed. The paper offers no domain-specific account of risks, suitable policies, or outcome measures for these activities.
- The recommended policies are not compared or stress-tested. Disclosure, mandatory human responsibility, automated triage, and greater emphasis on digestion are proposed without assessing compliance incentives, administrative costs, unintended consequences, or enforcement mechanisms.
- The paper does not address adversarial or strategic behavior. Researchers, institutions, journals, and AI vendors may optimize disclosures, benchmarks, citations, or exposition to appear compliant without improving correctness, understanding, or collective mathematical progress.
- The distribution of responsibility for AI-generated errors is unresolved. The paper does not specify how liability should be allocated among authors, tool developers, institutions, journals, referees, and publishers when an apparently verified result later proves defective.
- The long-term effects on the mathematical literature are unknown. It remains unclear whether AI-generated work will increase duplication, homogenize exposition, amplify historical biases, accelerate canonicalization, or make the literature less navigable.
- There is no empirical framework for measuring the broader goals of mathematics. Concepts such as community formation, aesthetic value, understanding, training, applicability, and contribution to cumulative knowledge remain listed but are not translated into indicators that could guide institutional decisions.
- The paper does not specify how competing mathematical values should be balanced. Open questions remain about trade-offs between speed and depth, accessibility and technical precision, formal rigor and conceptual insight, individual credit and collective production, and productivity and researcher development.
Practical Applications
Immediate Applications
The paper is primarily a normative and strategic essay rather than an empirical study. Its immediate applications therefore consist mainly of deployable changes to research workflows, publication practices, education, and institutional governance. Most depend on the paper’s Working Hypothesis that AI will soon perform a meaningful fraction of research-level mathematical tasks.
- AI-use disclosure in mathematical publications — Academia and scholarly publishing
- Journals, conferences, and preprint servers can add a mandatory Tool and computational resource disclosure section covering LLMs, proof assistants, automated theorem provers, literature-search systems, code-generation tools, and computational resources.
- Authors can specify which parts of a paper were AI-assisted, such as conjecture generation, proof search, formalization, editing, diagram production, or literature review.
- Category: Immediate Application.
- Potential tools/workflows: Standard disclosure templates, machine-readable metadata, reproducibility checklists, and journal submission fields.
- Dependencies and assumptions: Requires community agreement on what constitutes material AI assistance and policies that assign authors—not AI systems—responsibility for correctness, attribution, and interpretation.
- AI-assisted triage of mathematical manuscripts — Scholarly publishing
- Journals can use automated systems to flag missing citations, possible plagiarism, unverifiable claims, malformed formal proofs, inconsistent notation, inadequate explanations, or potentially fabricated references before assigning papers to human referees.
- This would conserve scarce expert reviewing capacity without eliminating peer review.
- Category: Immediate Application.
- Potential products: Submission-screening systems integrated with journal platforms; citation and attribution checkers; Lean, Rocq, or HOL compilation checks; exposition-quality diagnostics.
- Dependencies and assumptions: Automated flags must remain advisory. False positives, field-specific stylistic differences, confidentiality requirements, and bias in evaluation models would need monitoring.
- Formal proof verification as a routine research step — Mathematics and software
- Researchers can translate suitable arguments into proof-assistant languages such as Lean, Rocq, or HOL, allowing machine verification of logical correctness independently of author reputation.
- AI systems can assist with autoformalization, lemma discovery, proof completion, and conversion of informal arguments into formally checkable artifacts.
- Category: Immediate Application.
- Potential workflows: A paper repository containing both an ordinary exposition and a compilable formal proof; continuous integration for mathematical libraries; formal verification badges for journal articles.
- Dependencies and assumptions: Formalization remains costly for many areas, existing libraries may be incomplete, and a formally verified proof may still be difficult for humans to understand or use.
- Structured databases for open problems and partial results — Academia and research management
- Existing problem databases can be expanded to record conjectures, attempted approaches, failed strategies, partial results, formal proofs, provenance, literature links, and community verification status.
- AI systems could help classify submissions, identify related work, suggest known techniques, and detect duplicate or dependent claims.
- Category: Immediate Application.
- Potential products: Versioned open-problem registries, provenance graphs, formalized conjecture databases, and searchable repositories of negative results.
- Dependencies and assumptions: Databases require expert moderation, transparent licensing, stable identifiers, and mechanisms to prevent the rapid accumulation of low-quality AI-generated submissions.
- Transparent AI-assisted workflows for individual mathematicians — Academia and daily professional practice
- Researchers can use AI for literature search, notation conversion, code and diagram generation, drafting, proof exploration, and routine formalization while maintaining human review of every substantive claim.
- A practical workflow would preserve prompts, model versions, generated proof attempts, computational settings, and human revisions in an internal research log.
- Category: Immediate Application.
- Dependencies and assumptions: Tools must provide sufficient reliability, privacy, and access to relevant mathematical context. Researchers must not treat fluent output as evidence of correctness or originality.
- Human-centered standards for authorship and attribution — Academia, law, and research policy
- Institutions and publishers can explicitly state that authorship, credit, priority, and accountability belong to human contributors. AI systems should be acknowledged as tools rather than listed as authors.
- Authors can conduct additional attribution searches because AI systems may omit or misidentify intellectual sources.
- Category: Immediate Application.
- Potential tools: Provenance tracking, citation-recommendation systems, source-comparison tools, and contribution-taxonomy forms.
- Dependencies and assumptions: Attribution tools cannot guarantee complete historical credit; disciplinary norms and intellectual-property rules will continue to evolve.
- Assessment reforms in mathematics education — Education
- Instructors can redesign assignments so that learning is evaluated through oral examinations, handwritten or in-class reasoning, proof explanation, error analysis, reflective accounts of failed approaches, and live problem-solving rather than only final answers.
- AI can be used as a tutor or source of counterexamples, provided that students must critique, verify, and explain its output.
- Category: Immediate Application.
- Potential workflows: AI-generated hints followed by student validation; oral defenses of submitted proofs; assignments requiring comparison of multiple solution strategies.
- Dependencies and assumptions: Effective implementation requires instructor training, equitable access to tools, accommodations for students, and assessment methods that measure understanding rather than merely resistance to automation.
- Recognition of exposition, reviewing, and mathematical service — Academic institutions
- Departments, funders, and journals can give explicit credit for refereeing, textbook writing, formal-library maintenance, canonicalization of results, mentoring, and production of high-quality expository material.
- This directly addresses the paper’s concern that institutions over-reward proof generation and priority while under-rewarding the work that converts results into shared knowledge.
- Category: Immediate Application.
- Dependencies and assumptions: Promotion and funding systems must adopt broader evaluation criteria and develop credible ways to assess the quality and influence of community-maintenance work.
- Public communication about AI-generated mathematics — Policy and public outreach
- Mathematicians, professional societies, and science journalists can explain the distinction between generated, formally verified, human-understood, peer-reviewed, and canonicalized results.
- Public claims about AI achievements can include disclosure of the number of attempts, human scaffolding, computational cost, benchmark contamination, and evaluation procedures.
- Category: Immediate Application.
- Dependencies and assumptions: Requires access to reliable evaluation data and cooperation among researchers, AI developers, publishers, and journalists.
Long-Term Applications
The paper also implies broader institutional and technical transformations. These applications require further research, large-scale infrastructure, or sustained changes in mathematical culture because they concern the transition from proof scarcity to proof abundance.
- An end-to-end mathematical knowledge pipeline — Academia and research infrastructure
- A future platform could connect open problems to AI-generated candidate solutions, formal verification, human-readable exposition, peer review, community discussion, and eventual incorporation into textbooks or reference libraries.
- Each result could carry status labels such as unverified, formally verified, human-checked, peer-reviewed, community-accepted, and canonicalized.
- Category: Long-Term Application.
- Potential products: Research-management platforms, provenance-aware proof graphs, formal publication repositories, and APIs linking databases, proof assistants, journals, and educational resources.
- Dependencies and assumptions: Requires interoperable standards, persistent infrastructure, expert moderation, sustainable funding, and agreement on what qualifies as community acceptance or canonicalization.
- AI systems optimized for mathematical understanding rather than output volume — AI research and mathematics
- New models could be trained or evaluated not only on whether they produce correct proofs, but also on whether they identify the key idea, explain difficult steps, situate results in the literature, expose limitations, and generate useful generalizations.
- Evaluation could include expert teaching tests, proof-compression quality, transfer to related problems, and the ability of humans to reconstruct the argument.
- Category: Long-Term Application.
- Dependencies and assumptions: “Understanding” is difficult to operationalize. Evaluation must avoid creating another narrow benchmark that becomes vulnerable to Goodhart’s law.
- Human-readable formal mathematics and adaptive proof exposition — Education and software
- Proof assistants could eventually generate multiple synchronized representations of the same theorem: a machine-checkable proof, a concise expert proof, an introductory explanation, a visual dependency graph, and an interactive tutorial.
- Systems could preserve “natural friction” by highlighting the difficult or conceptually important steps rather than presenting every step as equally routine.
- Category: Long-Term Application.
- Potential products: Interactive textbooks, adaptive proof browsers, theorem dependency visualizers, and AI tutors grounded in formal libraries.
- Dependencies and assumptions: Requires advances in mathematical pedagogy, formal libraries, user modeling, and methods for distinguishing essential insight from routine elaboration.
- Large-scale collaborative formalization projects — Academia and open-source software
- Research communities could formalize major theories collaboratively, creating reliable libraries that serve simultaneously as verification infrastructure, educational resources, and training data for mathematical AI.
- Formalized theories could be reused in mathematics, computer science, cryptography, software verification, and safety-critical engineering.
- Category: Long-Term Application.
- Dependencies and assumptions: Formalization labor is substantial; projects need funding, contributor recognition, stable maintainers, shared coding standards, and tools capable of handling specialized or highly abstract mathematics.
- New publication venues for exposition, negative results, and partial progress — Scholarly communication
- The paper’s proof pipeline suggests a need for venues that publish verified but incomplete arguments, failed approaches, explanatory essays, formal proof artifacts, research talks, and canonical surveys.
- Such venues would reduce duplicated effort and provide a place for results that are valuable for collective understanding but do not constitute a conventional theorem paper.
- Category: Long-Term Application.
- Potential products: Peer-reviewed video journals, formal-proof supplements, negative-result repositories, and community-curated theory maps.
- Dependencies and assumptions: These venues require sustainable editorial models and changes in hiring, promotion, and grant evaluation so that researchers are rewarded for contributing to them.
- New priority and credit systems for an era of proof abundance — Research policy
- If AI generates many correct solutions, “first proof” may become a less useful measure of contribution. Institutions could instead recognize problem formulation, conceptual simplification, independent verification, exposition, attribution, generalization, software-library construction, and canonicalization.
- Contribution records could distinguish the originator of a conjecture, proof designer, formalizer, verifier, expositor, and curator.
- Category: Long-Term Application.
- Dependencies and assumptions: Priority disputes, fragmented contributions, and automated assistance complicate credit allocation. Any metric-based system would itself be vulnerable to Goodhart effects and should therefore use qualitative expert judgment.
- AI-supported peer-review cooperatives — Scholarly publishing and professional organizations
- Mathematical societies could create shared review infrastructures in which AI performs preliminary checking and experts receive credit for deep verification, explanation, and synthesis.
- Review assignments could be matched to subject expertise, workload, conflicts of interest, and the novelty or risk profile of a submission.
- Category: Long-Term Application.
- Potential products: Federated referee networks, reviewer-credit registries, confidential proof-analysis environments, and AI systems that produce structured review dossiers.
- Dependencies and assumptions: Confidentiality, reviewer incentives, liability, model auditing, and resistance to replacing expert judgment are central constraints.
- Canonical mathematical knowledge bases for downstream science and engineering — Software, healthcare, finance, robotics, and energy
- Once theories and proofs are formalized, explained, and accepted, they could be embedded more reliably into domain-specific computational systems.
- Examples include formally checked optimization algorithms for energy systems, verified numerical methods in engineering and climate modeling, certified statistical procedures in healthcare and finance, and theorem-backed control or planning algorithms in robotics.
- Category: Long-Term Application.
- Dependencies and assumptions: These applications depend on the mathematics becoming canonical and computationally implementable, as well as on domain validation, regulatory approval, numerical robustness, data quality, and integration with existing software systems. The paper does not itself establish efficacy in these sectors; it identifies the knowledge-infrastructure pathway that could enable them.
- Public policy for responsible AI-assisted mathematics — Government and international organizations
- Policymakers and mathematical societies could establish standards for disclosure, reproducibility, attribution, evaluation transparency, model access, archival preservation, and responsible use in publicly funded research.
- Funding programs could support open formal libraries, benchmark datasets containing genuinely novel problems, independent capability evaluations, and infrastructure for human review.
- Category: Long-Term Application.
- Dependencies and assumptions: Policies must remain adaptable as model capabilities change and should avoid privileging proprietary vendors or imposing compliance costs that disadvantage smaller institutions.
- A new division of labor between humans and mathematical AI — Daily life and professional knowledge work
- In the longer term, AI systems could handle routine derivations, symbolic manipulation, formal checking, and search across large mathematical libraries, while humans focus on choosing meaningful questions, interpreting results, communicating insight, and deciding which knowledge deserves adoption.
- This could improve access to advanced mathematical assistance for scientists, engineers, educators, and technically skilled members of the public.
- Category: Long-Term Application.
- Dependencies and assumptions: Broad benefits depend on trustworthy systems, affordable access, strong mathematical education, protection against opaque or incorrect advice, and preservation of human responsibility for consequential decisions.
Glossary
- Autoformalization: The automated translation of informal mathematical statements or proofs into a formal language suitable for machine verification. “Advances in AI, and in autoformalization into proof assistant languages”
- Canonicalization: The process of incorporating a mathematical result into the standard, definitive body of knowledge of a field. “This process of canonicalization”
- Compute: Computational resources expended to perform a task, especially by an AI system. “the compute expended”
- Conjecture: A mathematical proposition believed to be true but not yet proved. “not as a single conjecture, but as a family of conjectures indexed by a large number of free parameters.”
- Contamination: The influence of previously available information about a problem on the evaluation of an AI system’s ability to solve it independently. “the degree of contamination of the problem with prior literature”
- Definitive theory: The established, standardized theoretical framework into which accepted results are incorporated. “incorporated into the definitive theory of the field.”
- Erdős problems: A collection of mathematical problems posed by the mathematician Paul Erdős, often catalogued in an online database. “Sites devoted to collecting mathematical problems, such as the Erdős problems database”
- Formal proof: A proof expressed in a precisely defined formal system so that its correctness can be checked mechanically. “providing formal proofs where feasible and appropriate.”
- Formal system: A mathematical framework consisting of formal symbols, axioms, and rules of inference. “a formal system could not simultaneously be consistent, sufficiently expressive, and capable of proving its own consistency.”
- Formalization: The representation of mathematical concepts and arguments within a rigorous formal language or foundational system. “For further discussion of recent developments in formalization”
- Foundational framework: The formal system of basic concepts, axioms, and reasoning rules underlying mathematics. “an explicit, rigorous, and standardized foundational framework”
- Foundations of mathematics: The study of the fundamental logical and structural basis of mathematical reasoning. “the crisis in foundations”
- Frontier AI models: Advanced AI systems representing the current leading edge of capability. “an independent assessment of the capabilities of frontier AI models and harnesses”
- Generative AI: AI designed to produce new content, such as text, proofs, or diagrams, in response to input. “generative AI is inherently ungrounded”
- Gödel incompleteness theorems: Results showing that sufficiently expressive consistent formal systems contain true statements they cannot prove and cannot prove their own consistency. “the Gödel incompleteness theorems in 1931”
- Goodhart’s law: The principle that a measurement ceases to be a reliable indicator when it is made the target of optimization. “When a measure becomes a target, it ceases to be a good measure.”
- Harness: A software framework or collection of tools used to operate, evaluate, or augment an AI model. “four AI systems, using models publicly accessible as of May 28, 2026”
- Impedance mismatch: A structural incompatibility between the rate or capacity of different stages in a process. “significant ‘impedance mismatches’”
- Incompleteness: The property of a formal system in which some statements expressible in the system cannot be proved within it. “the Gödel incompleteness theorems”
- Machine learning system: A computational system that learns patterns or decision rules from data rather than being explicitly programmed for every case. “LLMs, machine learning systems, proof assistants, and other mathematical software.”
- Mathematical formalization: The conversion of informal mathematical reasoning into a formally specified representation that can be mechanically checked. “collaborative formalization projects”
- Mathematical scaffolding: Human assistance supplied to support or guide an AI system’s mathematical performance. “the amount of human scaffolding”
- Metamathematical: Relating to the study of mathematical systems, proofs, and statements from outside those systems. “It is a metamathematical one”
- Naive set theory: An informal approach to treating arbitrary collections as sets without sufficiently restrictive axioms. “The axioms of naive set theory contradicted each other”
- Orthogonal complement: A component conceptually independent of, and separated from, another component or question. “the ‘orthogonal complement’ of the AI Capability Conjecture”
- Peer review: The evaluation of scholarly work by qualified experts before publication. “overwhelm a traditional peer review system”
- Proof assistant: A computer program that enables users to construct and mechanically verify formal mathematical proofs. “Advances in AI, and in autoformalization into proof assistant languages”
- Proof digestion: The process by which a mathematical community understands, evaluates, disseminates, and incorporates a proof. “we need to decrease the emphasis that our culture places on proof generation”
- Proof exposition: The clear presentation and explanation of a mathematical proof for human readers. “Current AI tools have a decidedly mixed record with proof exposition.”
- Proof generation: The construction of a mathematical argument intended to establish the truth of a proposition. “Open problems” and “proof generation”
- Proof indigestion: The inability of mathematical institutions or communities to process an excessive volume of proofs. “proof indigestion”
- Proof verification: The process of checking whether a proposed proof is logically correct. “verify them to be correct”
- Proxy: A substitute measure used to represent or approximate another goal or property. “one could use one or two of these goals as convenient proxies for the others”
- Rocq: A formal proof assistant and programming language used to specify and verify mathematical reasoning. “such as Rocq, HOL, or Lean”
- Russell’s paradox: A contradiction arising from considering the set of all sets that do not contain themselves. “discoveries such as Russell's paradox in 1901”
- Set theory: The mathematical study of sets and their relationships, commonly used as a foundation for mathematics. “Practicing mathematicians proved theorems about sets, numbers, and infinities”
- Tacit knowledge: Practical or conceptual knowledge that is difficult to state explicitly and is transmitted through experience. “the tacit knowledge of a field is transmitted.”
- Theorem economy: A research culture that evaluates progress primarily by the production of new theorems. “the ‘theorem economy’ based primarily on proof generation.”
- Ungrounded: Lacking a direct guarantee that generated output corresponds to the underlying truth or intended property. “generative AI is inherently ungrounded”
- Verification: The systematic determination that a result or proof satisfies specified correctness conditions. “a proof whose correctness no longer depends on the reputation or the diligence of its author.”
- Working Hypothesis: A provisional assumption adopted to enable conditional analysis. “A reasonably strong version of the AI Capability Conjecture is true”
- Zero-shot evaluation: The assessment of a system on tasks it has not specifically been trained or fine-tuned to solve; in the paper, related assessment is described through novel problems. “genuinely novel mathematics”
