Papers
Topics
Authors
Recent
Search
2000 character limit reached

Training Mathematicians in the Age of AI: Intellectual Agency, Cognitive Offloading, and the PhD Thesis

Published 23 Sep 2026 in math.HO | (2609.28140v1)

Abstract: Powerful artificial intelligence is weakening the traditional relationship between mathematical output and evidence of mathematical expertise. In particular, the production of an original theorem or a polished dissertation can no longer, by itself, certify the intellectual formation of its nominal author. I argue that graduate mathematical education should therefore be organized around the formation of intellectual agency: internal technical competence, mathematical judgment, understanding, and responsible participation in a shared intellectual culture. I distinguish productive from premature cognitive offloading, propose complementary independent and AI-augmented modes of training, and suggest a corresponding reformulation of the role of the PhD thesis and dissertation defense. More broadly, I argue that academic mathematics should understand itself increasingly as an institution for the reproduction and stewardship of human mathematical expertise rather than primarily as a mechanism for producing theorems.

Authors (1)

Summary

  • The paper argues that traditional measures of mathematical expertise are corrupted by AI-generated proof, urging a focus on 'intellectual agency'—a set of skills centers on research direction, ethics integration, problem-solving evaluations, and interdisciplinary connections—that captures more typical achievements of graduate-level training.
  • Cognitive offloading should guide AI integration, distinguishing between 'productive offloading' (which enhances learning) and 'premature offloading' (which hinders skill development), recognizing a need for empirical research to determine this.

The paper argues that artificial intelligence has altered the epistemic relationship between mathematical output and mathematical expertise. A theorem, proof, or polished dissertation previously provided substantial evidence that its named author possessed technical competence, persistence, independence, and mathematical judgment. As AI systems increasingly generate sophisticated mathematical arguments, that inference becomes unreliable. The central proposal is therefore not to preserve a pre-AI conception of mathematical independence, but to reorganize graduate education around intellectual agency: the capacity to formulate worthwhile questions, evaluate arguments, direct research, integrate external assistance, explain mathematical significance, and participate responsibly in a shared intellectual culture.

The argument is normative rather than empirical. The paper reports no experiments, learning outcomes, or comparative assessment data, and it does not claim to provide a finalized institutional policy. Instead, it develops a conceptual framework for graduate training, cognitive offloading, dissertation assessment, and the social reproduction of mathematical expertise.

The epistemic problem created by AI

The paper begins from six axiomatic claims. Human expertise is desirable because it provides epistemic redundancy and permits machine outputs to be interrogated, criticized, and overruled. Tool access does not guarantee competent tool use: generating an apparently sophisticated argument may require little expertise, while diagnosing a subtle error may require years of mathematical training. Mathematical culture is not reducible to a set of formally valid propositions, because it includes interpretation, exposition, historical continuity, communal criticism, aesthetic judgment, and intergenerational transmission. Expertise must consequently be deliberately reproduced through apprenticeship and intellectual culture.

The fifth axiom is the decisive one: AI has permanently changed the evidentiary meaning of mathematical artifacts. This conclusion does not depend on continued capability improvements. Once systems capable of substantial mathematical assistance exist, a reader can no longer infer from a finished proof how much of the underlying mathematical cognition was performed by the nominal author. Disclosure norms may improve provenance, but they cannot restore the former presumption that sophisticated output ordinarily constituted evidence of sophisticated human activity. This is a strong claim about academic epistemology: the problem is not merely undisclosed AI use, but the loss of reliable information that was previously inferred from the artifact itself.

The implication is that theorem production can no longer function as a sufficient basis for judging mathematical formation. Theorems remain mathematically valuable, but their existence increasingly tells us less about the mathematical capacities of the person presenting them. The paper therefore distinguishes the value of a mathematical object from the evidentiary value of that object as a credential.

The discussion of formal verification sharpens this distinction. A machine-checked proof establishes that a formal statement follows from specified assumptions within a formal system; it does not by itself establish why the argument works, which ideas are essential, how the result relates to existing mathematics, or what mathematical meaning it has. Formal correctness and human understanding are thus treated as non-equivalent dimensions of mathematical achievement.

From independent researcher to intellectual agent

The conventional objective of doctoral training is often described as producing an independent researcher. The paper argues that this formulation is both historically inaccurate and conceptually inadequate. Mathematical research has always depended on advisors, collaborators, seminars, referees, literature, institutional resources, and inherited formal systems. Absolute independence is therefore a misleading ideal even before AI.

The proposed replacement is the intellectual agent. Intellectual agency is mathematical judgment exercised from within mathematics. It includes selecting important questions, assessing the feasibility of approaches, recognizing misunderstanding, identifying invalid arguments, rejecting poor suggestions from humans or machines, synthesizing disparate ideas, seeking outside expertise, and situating results within a broader research program. It also includes post-proof activity: reformulating results, identifying conceptual structure, generating consequences, and determining what other mathematicians should learn from the work.

Technical proficiency is a precondition for this agency, but agency is not equivalent to technical virtuosity. A mathematician need not execute every difficult argument unaided. They must instead possess enough internal mathematical structure to recognize the nature of a problem, distinguish conceptual from technical obstacles, assess whether a generalization is natural, and decide whether an external argument merits trust. The paper accordingly rejects the view that AI-supported work is necessarily non-independent. Heavy reliance on AI can coexist with substantial agency, while solitary work and technically strong output can coexist with weak agency over question selection, interpretation, and synthesis.

This conception overlaps with mathematical judgment as formulated in the related graduate-training proposal by Glickenstein, which emphasizes proof validation, definition assessment, monitoring understanding, transfer, and tool governance (Glickenstein, 17 Sep 2026). The present paper adopts a broader frame: mathematical judgment is one component of intellectual agency, which additionally encompasses research direction, synthesis, recognition of when expertise is needed, and stewardship of disciplinary culture.

Cognitive offloading and the formation of competence

The paper’s treatment of cognitive offloading is its principal account of how AI should be integrated into graduate formation. It rejects both unrestricted delegation and blanket technological abstinence. The relevant distinction is between productive offloading and premature offloading.

Productive offloading removes unnecessary drudgery, increases the scale of exploration, or permits attention to be redirected without undermining the competence required for judgment. Premature offloading delegates a task whose performance is itself part of the process by which the relevant competence would have formed. The distinction is developmental rather than moral: the same task may be productively delegated by an established expert but prematurely delegated by a student who has not yet internalized the underlying skill.

This framework yields two complementary modes of training. In independent mode, students work without AI assistance when the pedagogical objective is to develop or assess capacities that AI could otherwise perform, including sustained reasoning, reconstruction of details, management of failed approaches, and detection of invalid arguments. In augmented mode, students use AI extensively and deliberately in research, exploration, verification, exposition, and synthesis.

The independent mode is explicitly described as diagnostic and developmental, not ethically superior. AI is not treated as a contaminant of mathematical practice. Rather, an assessment cannot provide evidence that a student has internalized a capacity if the assessment permits an external system to perform that capacity. This point aligns with the distinction between AI-independent proving and AI-assisted proving in mathematics education, but the paper extends the distinction from individual assignments to the full process of doctoral formation (2609.28140).

The paper also identifies an institutional tension. Graduate education and academic employment continue to reward papers containing original theorems, while AI increasingly makes theorem production a poor proxy for human mathematical formation. Tao’s discussion of a transition from “proof scarcity” to “proof abundance” is used to characterize this change (Tao, 17 Aug 2026). The argument is not that theorems have become unimportant. Rather, an incentive system that treats theorem production as the principal currency of credit and prestige is no longer adequate if institutions intend to cultivate human mathematical expertise.

The paper leaves the operational boundary between productive and premature offloading unresolved. It concedes that this boundary will vary by individual, mathematical field, career stage, and task, and that empirical evidence will be needed. This concession is important: the conceptual distinction is clear, but the paper does not offer a validated protocol for determining which tasks students should perform unaided at particular stages of training.

Mathematics as an intellectual network

The paper situates individual mathematical achievement within a large network of dependencies. Definitions, proof techniques, terminology, foundational results, pedagogical practices, and research questions are inherited from prior generations. Advisors, collaborators, seminar participants, referees, students, and institutions all contribute to the production and validation of mathematical work. Authorship has traditionally compressed this distributed infrastructure into a small number of names.

AI makes this networked character more visible because the individual artifact is no longer reliable evidence of the quantity or kind of human cognition behind it. The appropriate professional question is consequently not only what a mathematician proved, but what their intellectual agency helped the community understand. This reorientation elevates teaching, exposition, criticism, synthesis, organization, mentoring, and institutional maintenance from peripheral activities to central components of mathematical stewardship.

The paper is careful not to equate stewardship with conservation. A steward may criticize inherited ideas, reorganize a theory, discard established approaches, or redirect collective attention. Nor does networked production weaken the importance of attribution. It instead supports more accurate attribution by recognizing forms of intellectual and institutional labor that theorem-centered credit systems often obscure.

A significant risk follows from this proposal. Judgment, taste, agency, and stewardship are less legible than theorem counts and therefore more vulnerable to patronage, prestige effects, and arbitrary authority. Replacing crude quantitative metrics with opaque elite judgment would create a different but equally serious failure. The paper therefore calls for distributed evaluation, multiple forms of evidence, transparent criteria where possible, careful attribution, and resistance to any single advisor or eminent mathematician becoming the arbiter of mathematical value.

Reconceiving the dissertation and defense

The dissertation should no longer function principally as evidence that a candidate has produced an original theorem. It should instead provide evidence of mathematical formation and intellectual agency. Original mathematical propositions remain an important component, but they are placed within a broader evidentiary structure.

The proposed dissertation has several components:

  • A substantive mathematical contribution: The thesis should ordinarily contain significant mathematical work, whether or not AI participated in its production.
  • An intellectual narrative: The candidate should explain how the problem arose, which background is relevant, where the main ideas came from, and which approaches failed.
  • Synthesis: The thesis should locate the contribution within a wider mathematical landscape and explain what other mathematicians should understand or be able to do as a result.
  • Machine-assisted methodology: Where AI was used, the candidate should document which tasks were delegated, how outputs were checked or rejected, and where human judgment entered.
  • Independent reconstruction and defense: The candidate should be able to evaluate, reconstruct, modify, and defend relevant mathematics without delegating those capacities during the assessment.

The proposal assigns substantially greater evidentiary weight to the dissertation defense. Examiners might alter hypotheses, ask what breaks under modified assumptions, request reconstruction of a proof component, compare alternative approaches, identify conceptual landmarks, or ask the candidate to articulate plausible extensions. Such questioning tests understanding and agency rather than merely familiarity with a polished document. The paper’s proposal is closely related to “mathematical ownership” in Glickenstein’s account of AI-aware dissertation defenses (Glickenstein, 17 Sep 2026).

The defense is not presented as an infallible authenticity test. Oral performance is sensitive to language, disability, personality, and pressure, and live examination may reward speed or charisma rather than depth. It should therefore form one part of a broader evidentiary record. The central institutional principle is not that oral examination solves the provenance problem, but that assessment must include settings in which students demonstrate capacities that polished AI-assisted artifacts cannot reliably reveal.

Graduate education as apprenticeship and stewardship

The paper ultimately redefines doctoral training as an apprenticeship toward stewardship of mathematics and the formation of intellectual agency. Independent problem solving remains essential because it develops technical proficiency, but it is placed alongside activities that are often treated as secondary: research seminars, exposition, teaching, collaborative research, refereeing, synthetic writing, AI-assisted exploration, and mentoring.

These activities matter because they increase the community’s capacity to understand, criticize, use, and extend mathematics. In a high-productivity environment, conceptual synthesis, exposition, literature organization, formalization, and the construction of shared mathematical infrastructure may become more consequential than the production of another isolated theorem. The paper does not claim that these contributions should replace original research in every doctorate; it claims that they should receive substantially greater recognition in the training and evaluation of mathematicians.

The emphasis on collaboration follows from the same analysis. As AI expands the volume of potentially relevant mathematical output, human attention and coordination become scarce resources. Graduate students therefore need to learn how to combine human and machine-generated knowledge, communicate partial understanding, attribute credit, disagree productively, and recognize better judgment in collaborators.

The paper also addresses recruitment and professional identity. Students should not be given a nostalgic account in which mathematics consists primarily of writing papers containing proofs of finished results. At the same time, they should not be told that mathematical training has become meaningless. The paper’s claim is more specific: the content of mathematical work may change, but graduate education can remain valuable if it forms people capable of understanding, directing, evaluating, and stewarding mathematical activity under conditions of proof abundance.

Limitations and open questions

The paper’s principal limitation is its reliance on normative axioms rather than empirical validation. It assumes that preserving human mathematical expertise, epistemic self-government, and a living human mathematical culture are desirable. These assumptions are defended through arguments about robustness, intergenerational transmission, and the value of mathematical experience, but alternative institutional priorities are not systematically analyzed.

The proposal also leaves unresolved how intellectual agency should be measured. The paper correctly observes that agency and stewardship are difficult to operationalize, but it provides no rubric, reliability analysis, or evidence that defenses, portfolios, or distributed evaluation can distinguish genuine understanding from rehearsed performance. Its warning about charisma, patronage, and prestige effects applies directly to the proposed evaluative framework.

A further open question concerns the developmental schedule for independent and augmented modes. The paper does not specify which mathematical tasks should be AI-free, for how long, or according to what criteria a student may transition from restricted to extensive assistance. Nor does it determine whether competence acquired through AI-augmented pathways is equivalent to competence acquired through traditional independent work. These are empirical questions about learning, transfer, retention, and judgment.

Finally, the essay assumes that academia remains sufficiently adaptable to serve as the principal institution for reproducing mathematical expertise. It acknowledges institutional weaknesses and the possibility of new forms of organization, but it does not examine how existing incentives, labor conditions, unequal access to AI, or disciplinary hierarchies might obstruct the proposed reforms. The claim that AI should broaden recognized mathematical contribution therefore remains dependent on substantial changes in evaluation and resource allocation.

Conclusion

The paper’s central thesis is that AI has weakened the dissertation and theorem as proxies for mathematical expertise. Graduate education should consequently certify not only mathematical output but also internal competence, judgment, understanding, intellectual agency, and stewardship of mathematical culture. Independent and AI-augmented modes of training should be used for complementary purposes, while dissertations and defenses should provide evidence of both mathematical contribution and responsible cognitive participation. The resulting ideal is not the isolated independent researcher, but the technically competent intellectual agent who can direct, evaluate, explain, and sustain mathematical work in a networked and AI-assisted discipline.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is the paper about?

This paper asks an important question:

How should mathematicians be trained now that artificial intelligence can help create difficult mathematical proofs and even discover new theorems?

In the past, a strong thesis or a difficult proof was good evidence that a student understood mathematics deeply and had done the work independently. The paper argues that this is no longer always true. AI can now produce impressive mathematical work, so a finished paper does not necessarily show how much the human author understands.

The author’s main idea is that mathematics education should focus less on producing theorems alone and more on developing intellectual agency. This means being able to understand mathematics, make good decisions, judge whether an argument is correct, ask worthwhile questions, and use tools responsibly.

2. What questions does the paper explore?

The paper focuses on several connected questions:

  • What should a mathematics PhD prove about a student in the age of AI?
  • What abilities should graduate students develop before relying heavily on AI?
  • When does using AI save time in a helpful way, and when does it prevent learning?
  • How can teachers tell whether a student truly understands work produced with AI?
  • Should a dissertation be judged mainly by its new theorems?
  • What role should mathematicians continue to play if AI can solve more and more problems?

The author argues that a mathematician should not be seen as someone who works completely alone. Instead, a mathematician should be an intellectual agent: someone who can direct, understand, question, improve, and explain mathematical work, whether or not machines are involved.

3. How does the paper approach these questions?

This is mainly a philosophical and educational essay, not an experiment with data or a mathematical proof. The author develops his argument through several ideas, or “axioms,” which are basic principles he believes should guide mathematics education.

Human expertise is still necessary

The paper argues that society needs human experts who can check and question AI. The author compares this to airplanes. Autopilot can perform many tasks, but trained pilots are still needed when something unexpected happens.

In the same way, mathematicians should be able to:

  • Check whether an AI-generated argument makes sense
  • Find mistakes
  • Explain why a proof works
  • Decide whether a problem is important
  • Reject bad suggestions
  • Take responsibility for the final result

An AI may be good at producing an argument, while a human may need years of training to judge whether that argument is actually correct.

Correctness is not the same as understanding

The paper also distinguishes between a proof being formally correct and a person understanding it.

For example, a computer program such as Lean can check that a proof follows logically from certain assumptions. This is useful, but it does not automatically explain:

  • Why the result matters
  • What the main idea of the proof is
  • How the result connects to other mathematics
  • What new questions the result suggests

A computer spell-checker can tell you whether words are spelled correctly, but it cannot necessarily explain the meaning or beauty of a poem. The author says formal proof-checking is similar: it checks correctness, but it is not the same as human understanding.

Cognitive offloading

The paper uses the term cognitive offloading for letting a tool do part of our thinking. Using a calculator instead of doing long arithmetic by hand is one example.

The author divides this into two types:

  • Productive offloading: Using a tool to avoid boring work without losing the ability to understand the important ideas.
  • Premature offloading: Letting a tool do a task before we have learned how to do it ourselves.

For example, an experienced mathematician might use AI to test examples or organize calculations. This could save time. But a beginner who asks AI to solve every problem may never develop the ability to reason independently.

The paper suggests two training modes:

Training mode Main purpose
Independent mode Students work without AI when they need to develop or demonstrate skills such as reasoning, proving, and checking arguments.
Augmented mode Students use AI to explore ideas, test possibilities, and carry out research, while still explaining and checking the results.

The goal is not to ban AI. It is to make sure students learn which abilities must remain inside their own minds.

Mathematics as a shared culture

The paper also says that mathematics is not just a pile of correct statements. It is a human activity involving:

  • Teaching and learning
  • Asking questions
  • Discussing ideas
  • Writing explanations
  • Correcting mistakes
  • Sharing discoveries
  • Building on earlier generations

A theorem usually depends on many people: teachers, advisors, authors of earlier books, seminar audiences, collaborators, referees, and students. Therefore, mathematicians should not only produce results. They should also help preserve and improve the shared culture of mathematics.

4. What are the main findings or conclusions?

Because this is an argumentative essay rather than an experimental study, its “findings” are mainly recommendations and conclusions.

A theorem alone is no longer enough evidence

In the past, a difficult theorem or polished dissertation strongly suggested that its author had developed significant mathematical skill. AI weakens this connection because a machine may have produced much of the work.

This does not mean the theorem is unimportant. It means that the theorem alone tells us less about the human who presents it.

Mathematical training should develop intellectual agency

The author says students should learn to:

  • Understand the mathematics they use
  • Choose important and interesting questions
  • Judge whether an argument is reliable
  • Recognize mistakes and weak ideas
  • Explain the main concepts clearly
  • Connect new work to existing knowledge
  • Decide when to ask other people or machines for help
  • Use AI honestly and responsibly

Technical skill is still important. A student cannot judge advanced mathematics without knowing enough mathematics themselves. However, students do not need to perform every tedious calculation without tools. They need enough internal understanding to know what the calculations mean and whether the answer can be trusted.

The dissertation should show mathematical formation

The paper proposes that a PhD dissertation should include more than an original result. It should also show:

  1. A mathematical contribution The student should normally make a real contribution to mathematics.
  2. An intellectual story The thesis should explain how the problem arose, what background is needed, and which approaches succeeded or failed.
  3. Synthesis The student should explain how the work fits into the larger field and what other mathematicians can learn from it.
  4. Honest discussion of AI use If AI was used, the student should explain what it did, how its suggestions were checked, and where human decisions were important.
  5. Strong defense and questioning During the defense, the student should be able to explain the work, change assumptions, identify limitations, reconstruct parts of arguments, and discuss future research.

The defense should not be treated as a perfect test. Some people are nervous or communicate differently in oral exams. Instead, it should be one part of a larger collection of evidence about the student’s understanding.

More kinds of mathematical work should be valued

If AI makes theorem production much easier, universities may need to recognize other important contributions, such as:

  • Clear explanations of difficult ideas
  • Organizing confusing research literature
  • Creating useful examples or frameworks
  • Teaching and mentoring
  • Checking and criticizing arguments
  • Connecting different areas of mathematics
  • Helping other mathematicians understand and use new results
  • Building systems for formal proof and verification

The paper suggests that these activities may become even more valuable when AI produces mathematical information faster than humans can understand it.

5. Why is this research important?

The paper’s main warning is that universities should not continue judging mathematicians mainly by the number of theorems or papers they produce. If AI can create many correct proofs, then counting results may no longer show who has genuine expertise.

Instead, mathematics education should help create people who can:

  • Think carefully
  • Understand difficult ideas
  • Make responsible choices
  • Work with both humans and machines
  • Explain knowledge to others
  • Notice errors
  • Preserve and improve mathematical culture

The author is not arguing that AI will destroy mathematics. In fact, AI could help mathematicians explore more ideas, test more possibilities, and make new discoveries. But people will need enough knowledge to guide AI and understand its results.

The paper’s final message is simple:

AI may change what mathematicians do, but it should not remove mathematicians from mathematics.

A future mathematician may use AI often, but should still be able to understand, question, explain, and take responsibility for mathematical work. The purpose of a PhD should therefore be to show that a person has become a capable and responsible participant in the mathematical community—not merely that a theorem has appeared under their name.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

The paper offers a normative framework for graduate mathematics education in an AI-rich environment, but leaves the following issues unresolved:

  • Lack of empirical evidence: The claims that AI weakens the evidentiary value of dissertations, increases the gap between mathematical generation and evaluation, and affects cognitive development are not supported by systematic studies of students, mathematicians, or existing AI-assisted research practices.
  • Unclear definition of intellectual agency: The paper identifies intellectual agency through capacities such as judgment, synthesis, question selection, and stewardship, but does not provide measurable criteria for determining when a student has acquired these capacities.
  • No operational test for technical proficiency: The paper argues that students need substantial internal competence but does not specify the minimum knowledge, skills, or field-specific standards required for meaningful agency.
  • Undefined boundary between productive and premature offloading: The distinction is conceptually useful but lacks practical rules for deciding which tasks may safely be delegated at different stages of training or across different mathematical subfields.
  • Insufficient account of disciplinary variation: The proposed training and assessment model is not adapted to differences between pure and applied mathematics, theoretical computer science, statistics, formal mathematics, computational fields, or highly collaborative research areas.
  • Unresolved trade-off between independent and augmented modes: The paper does not determine how much time students should spend in AI-free versus AI-assisted work, how these modes should be sequenced, or whether early restrictions might disadvantage students who learn most effectively through interactive tools.
  • Unvalidated assessment methods: The proposed defenses, oral examinations, reconstruction tasks, and synthesis components have not been shown to measure understanding or agency reliably, consistently, or better than current assessments.
  • Risk of assessment bias: Although the paper acknowledges that oral examinations may disadvantage students because of personality, language, disability, or performance anxiety, it does not develop concrete accommodations or alternative assessment formats that preserve validity.
  • Potential subjectivity of qualitative criteria: Concepts such as “taste,” “judgment,” “conceptual understanding,” and “stewardship” remain difficult to evaluate transparently and could reproduce existing status hierarchies, disciplinary fashions, or advisor biases.
  • No framework for examiner calibration: The paper does not explain how committees would be trained to assess intellectual agency consistently or how disagreements among examiners would be resolved.
  • Unresolved provenance and disclosure standards: The proposed disclosure of AI assistance is not specified in enough detail to determine what uses must be reported, how prompts and model outputs should be documented, or how disclosure requirements should vary by task.
  • Limits of transparency: The paper does not address whether AI-generated reasoning can be meaningfully documented when models produce opaque, nondeterministic, or rapidly changing outputs, nor how institutions could verify a candidate’s account of tool use.
  • No analysis of false positives and false negatives: The proposed procedures may wrongly classify capable students as lacking agency or accept students who have outsourced substantial cognition; the paper does not examine the reliability of these errors or their consequences.
  • Unclear relationship between human understanding and formal verification: The paper distinguishes machine-checkable correctness from understanding but does not specify how formal proof development, proof assistants, or computational experimentation should contribute to evaluating mathematical competence.
  • Insufficient treatment of collaborative authorship: The paper emphasizes networked mathematical production but does not explain how doctoral programs should assess individual agency within large collaborations involving advisors, research groups, software, and AI systems.
  • Unresolved status of AI as a participant in mathematical culture: The paper leaves open whether AI-generated conjectures, proofs, expositions, or critiques should count as contributions to mathematical knowledge and how credit should be allocated among humans, institutions, and systems.
  • Limited attention to incentives and career structures: The paper calls for rewarding synthesis, exposition, mentoring, and stewardship but does not propose changes to hiring, promotion, grant, journal, or ranking systems that currently prioritize publications and theorems.
  • Possible unintended expansion of doctoral requirements: Adding independent work, AI-assisted research, teaching, mentoring, refereeing, synthesis, disclosure, and expanded defenses could lengthen doctoral training or increase workload without clarifying which existing requirements should be reduced.
  • Equity implications remain underdeveloped: The paper explicitly excludes broader inequalities from its scope, leaving unresolved how unequal access to advanced AI tools, computing resources, expert mentorship, and institutional support could affect admissions, training, and professional opportunities.
  • Accessibility concerns are not addressed in depth: The framework does not determine how AI restrictions should be adapted for students with disabilities or for students who rely on assistive technologies without undermining the validity of assessments.
  • Uncertain effects on student motivation and identity: The paper does not provide evidence about how shifting emphasis away from theorem production would affect students’ motivation, sense of achievement, career aspirations, or willingness to pursue mathematical research.
  • No account of AI capability thresholds: The recommendations assume “powerful” AI but do not specify how educational policies should change as systems move from assistance with routine tasks to reliable theorem proving, research direction, or autonomous mathematical discovery.
  • Institutional feasibility is not demonstrated: The proposed reforms would require additional faculty time, examiner expertise, mentoring capacity, and administrative coordination, but the paper does not assess whether departments can implement them under existing resource constraints.
  • Unresolved governance and enforcement questions: The paper does not explain who should set AI-use policies, how violations should be investigated, what sanctions are appropriate, or how policies can remain effective as tools change rapidly.
  • Limited consideration of non-academic mathematical careers: The framework is centered on academic mathematics and doctoral stewardship, leaving unclear how intellectual agency and AI-assisted competence should be developed or evaluated for students pursuing industry, government, education, or interdisciplinary careers.
  • No longitudinal evidence about skill retention: It remains unknown whether AI-free training produces durable technical competence or whether skills decline once graduates enter AI-intensive research environments.
  • Open question about the value of friction: The paper argues that some difficulty is educationally valuable and some is wasted effort, but does not identify how educators can distinguish the two or determine when removing a task improves rather than weakens learning.
  • Unresolved relationship between human expertise and superior AI performance: If AI systems eventually outperform humans in judgment, synthesis, and error detection, the paper does not explain which specifically human capacities should remain educational priorities or how their value should be justified.
  • No comparative evaluation of alternative institutional models: The paper assumes that academia is well positioned to adapt, but does not compare universities with independent institutes, online communities, professional societies, or AI-mediated learning environments as sites for reproducing mathematical expertise.

Practical Applications

The paper is primarily a normative and institutional proposal rather than an empirical study. Its practical implications therefore concern how universities, professional communities, and AI developers should redesign training, assessment, attribution, and knowledge workflows so that human mathematical expertise remains meaningful in an AI-rich environment.

Immediate Applications

  • Redesign doctoral assessment around intellectual agency (higher education; deployable now)
    • understanding of the relevant literature;
    • ability to explain why a problem matters;
    • distinction between conceptual ideas and technical machinery;
    • identification of failed approaches and limitations;
    • ability to modify hypotheses and predict what breaks;
    • explanation of future research directions.
    • Potential tool or workflow: a dissertation rubric that scores technical competence, judgment, synthesis, communication, AI-use transparency, and stewardship alongside originality.
    • Dependencies: faculty agreement on standards, examiner training, and safeguards against subjective judgments about “taste” or “agency.”
  • Introduce independent and AI-augmented training modes (mathematics education; deployable now)
    • independent mode: no generative AI when the goal is to build or measure foundational reasoning;
    • augmented mode: AI is permitted or encouraged when the goal is research exploration, auditing, comparison of approaches, or productivity.
    • This can be implemented in coursework, qualifying examinations, research rotations, and thesis milestones.
    • Dependencies: departments must define which competencies must first be internalized and avoid treating AI abstinence as a moral requirement rather than a pedagogical choice.
  • Reform qualifying and oral examinations (higher education; deployable now)
    • audit machine-generated proofs;
    • locate errors or missing assumptions;
    • reconstruct arguments independently;
    • compare multiple proposed solutions;
    • explain when an AI-generated conjecture is implausible.
    • Potential product: assessment platforms with configurable AI access, logging, and examiner-generated oral follow-ups.
    • Dependencies: accessibility accommodations, recognition that oral performance is imperfect, and multiple forms of evidence rather than reliance on a single examination.
  • Require transparent AI-use statements in theses and papers (academic publishing and research governance; deployable now) Authors can document which AI systems were used, for what tasks, and how outputs were checked. Disclosure should focus on intellectual provenance and reproducibility rather than attempting to detect AI use through unreliable forensic methods. Potential workflow: a standardized “AI contribution and verification” appendix recording prompts or task categories, machine-generated material retained or discarded, human verification procedures, and use of formal proof assistants. Dependencies: clear privacy rules, evolving disclosure standards, and distinction between editorial assistance, computational assistance, conjecture generation, and substantive mathematical reasoning.
  • Expand the role of the dissertation defense (higher education; deployable now)
    • changing assumptions in a central theorem;
    • explaining the intuition behind a proof;
    • reconstructing a key lemma without notes;
    • identifying dependencies on external tools;
    • comparing alternative proof strategies;
    • explaining what the student would investigate next.
    • Potential workflow: a defense protocol combining presentation, technical questioning, AI-use discussion, and a later written or supervised reconstruction task.
    • Dependencies: examiner consistency, accommodations for language and disability, and protection against confusing charisma or rapid recall with deep understanding.
  • Teach AI auditing and mathematical verification as explicit skills (education, software, and research; deployable now) Courses can train students to treat AI outputs as hypotheses requiring evaluation. Exercises could involve intentionally flawed AI-generated proofs, ambiguous definitions, invalid generalizations, and formally correct but conceptually uninformative arguments. Potential products: educational systems that generate candidate proofs and ask students to classify, repair, or explain them rather than merely accept them. Dependencies: reliable benchmarks of mathematical error types and instructors capable of distinguishing technical correction from genuine conceptual understanding.
  • Use formal proof systems as verification tools, not substitutes for understanding (formal methods and software; deployable now) Mathematicians can combine systems such as Lean with human-readable explanations. AI may propose formalizations or proof steps, while researchers remain responsible for selecting definitions, checking assumptions, and interpreting the result. Applications: verified libraries, machine-assisted proof debugging, curriculum modules in formal mathematics, and reproducible computational supplements to papers. Dependencies: formalization cost, availability of libraries, limitations of current proof assistants, and the paper’s distinction between formal correctness and conceptual comprehension.
  • Recognize exposition, synthesis, mentoring, and teaching as research-training activities (academia; deployable now)
    • organizing fragmented literature;
    • writing conceptual surveys;
    • maintaining mathematical libraries;
    • mentoring junior students;
    • leading seminars;
    • producing high-quality explanations of difficult theories;
    • refereeing and constructive criticism.
    • These activities directly develop the stewardship capacities emphasized in the paper.
    • Dependencies: changes to promotion and funding criteria, reliable attribution, and protection against assigning unpaid service disproportionately to junior or marginalized researchers.
  • Develop AI-use policies based on the learning objective, not blanket prohibition (policy and educational administration; deployable now) Institutions can classify assignments according to whether they assess:

    1. independent internal competence;
    2. responsible AI-assisted problem solving;
    3. auditing and verification;
    4. communication and synthesis. Policies can then specify permissible tools and required disclosure for each category. Dependencies: regular policy revision as capabilities change, student access to comparable tools, and attention to privacy, bias, and unequal computing resources.
  • Improve everyday mathematical and technical decision-making through verification habits (daily life and professional work; deployable now)

    • check assumptions;
    • test results against simple cases;
    • recognize implausible conclusions;
    • ask for alternative explanations;
    • avoid delegating the very reasoning they are trying to learn.
    • Dependencies: users need baseline domain knowledge, and AI interfaces should make uncertainty, assumptions, and provenance visible rather than presenting polished answers as authoritative.

Long-Term Applications

  • Create doctoral programs centered on intellectual agency and stewardship (higher education; requires institutional redesign)
    • a substantive mathematical contribution;
    • independent foundational work;
    • an AI-augmented research project;
    • a synthesis or exposition project;
    • teaching or mentoring;
    • evidence of collaboration and responsible attribution;
    • a defense demonstrating understanding and judgment.
    • Potential outcome: the PhD becomes certification of readiness to participate in and steward a mathematical community, not merely certification of theorem production.
    • Dependencies: consensus about standards, comparable evaluation across institutions, funding, and methods for preventing qualitative assessment from becoming opaque or elitist.
  • Develop longitudinal measures of mathematical agency (education research and academic policy; requires further research)
    • select worthwhile problems;
    • monitor their own understanding;
    • distinguish conceptual from technical difficulty;
    • detect machine-generated nonsense;
    • integrate external suggestions;
    • explain and extend results;
    • recognize when expert help is needed.
    • Potential product: multi-year assessment dashboards combining coursework, oral examinations, research artifacts, peer feedback, and reflective explanations.
    • Dependencies: empirical validation, avoidance of narrow proxies, protection against surveillance, and sensitivity to different mathematical fields and cultures.
  • Build AI systems designed to preserve rather than replace mathematical learning (AI, education technology, and software; requires development)
    • graduated hints;
    • Socratic questioning;
    • proof decomposition;
    • error localization without immediate correction;
    • requests for the learner’s explanation;
    • delayed solutions;
    • adaptive switching between independent and augmented modes.
    • Potential products: graduate-level mathematical tutors and research assistants that explicitly monitor premature cognitive offloading.
    • Dependencies: accurate modeling of learner competence, explainability, strong pedagogical evidence, and safeguards against the system misjudging what a student already understands.
  • Create provenance-aware mathematical research environments (research infrastructure and publishing; requires scaling)
    • human ideas and decisions;
    • AI-generated conjectures or proof attempts;
    • formal verification;
    • literature sources;
    • collaborators’ contributions;
    • revisions and rejected approaches.
    • Such systems could produce richer contribution records than a conventional author list.
    • Potential tools: version-controlled research notebooks, proof-development repositories, machine-readable contribution statements, and provenance graphs linking claims to evidence.
    • Dependencies: interoperability, intellectual-property policy, privacy, resistance to excessive monitoring, and recognition that provenance records cannot fully capture tacit understanding.
  • Reform academic incentives and research evaluation (science policy and higher education; long-term)
    • conceptual synthesis;
    • reliable exposition;
    • mentoring;
    • formalization and reproducibility;
    • research-community infrastructure;
    • correction of errors;
    • collaboration;
    • contributions that expand collective understanding.
    • The number of papers or theorems would remain relevant but would no longer function as the dominant proxy for expertise.
    • Dependencies: transparent criteria, distributed evaluation, protection from prestige hierarchies, and evidence that alternative metrics are less gameable than current ones.
  • Establish new professional roles for mathematical stewardship (academia, industry, and public policy; long-term) AI-rich organizations may need specialists who curate mathematical knowledge, audit automated reasoning, maintain formal libraries, assess model outputs, and translate between machine-generated mathematics and human decision-makers. Potential roles: mathematical auditor, formalization engineer, proof-systems curator, AI-assisted research coordinator, and mathematical knowledge architect. Dependencies: labor-market demand, recognized credentials, clear liability arrangements, and sufficient human expertise to supervise automated systems.
  • Apply the intellectual-agency model to other high-stakes sectors (healthcare, finance, engineering, law, and public administration; long-term)
    • clinicians auditing diagnostic suggestions;
    • engineers reviewing automated designs;
    • financial analysts challenging model assumptions;
    • policymakers evaluating algorithmic forecasts;
    • software engineers inspecting AI-generated code and proofs.
    • Dependencies: sector-specific competency standards, legal accountability, explainable systems, and access to independent experts capable of overruling automated recommendations.
  • Develop collaborative human–AI mathematical research networks (research and industry; long-term) As theorem and conjecture generation becomes more abundant, research communities could organize around human selection, interpretation, synthesis, and validation. AI systems might search vast spaces of conjectures, while human mathematicians determine which questions are meaningful, connect results across fields, and decide how findings should enter the discipline. Potential outcome: research workflows optimized for “proof abundance,” with human attention directed toward conceptual structure and collective understanding. Dependencies: scalable verification, methods for ranking significance rather than merely correctness, sustainable curation, and preservation of human participation rather than complete automation.
  • Preserve human mathematical culture as a public and educational good (cultural policy and daily life; long-term) Institutions could support seminars, mathematical communities, public explanation, historical preservation, and accessible educational resources even if automated systems can generate correct mathematics more cheaply. This follows the paper’s claim that mathematics includes curiosity, shared discovery, interpretation, and intergenerational transmission—not only formally true propositions. Dependencies: public and institutional willingness to fund activities whose value is cultural and epistemic rather than immediately commercial, as well as inclusive access to mathematical communities.

Glossary

  • Affective dimension: The role of emotions and subjective experiences in a practice or field. “There is also an affective dimension to mathematical practice that should not be treated as incidental.”
  • Apprenticeship: A structured process in which a novice develops expertise through guided participation with experienced practitioners. “Expertise is formed through an apprenticeship-like process in which one is exposed to examples, diverse viewpoints, books, conversations, and seminars.”
  • Cognitive authorship: The substantive ownership of the reasoning and intellectual work represented by a scholarly artifact. “Sathi's distinction between cognitive authorship and what he calls the institutional or ``administrative'' author is useful here.”
  • Cognitive offloading: Delegating mental tasks to external tools or systems in order to reduce cognitive effort. “Because technical proficiency is a precondition for intellectual agency, graduate education must distinguish forms of offloading that preserve or deepen proficiency from those that prevent proficiency from forming.”
  • Conceptual synthesis: The process of integrating separate ideas or results into a coherent conceptual understanding. “Novel viewpoints, excellent expositions of difficult theories, the organization of confused literatures, conceptual synthesis, formalization, and the creation of structures that make other mathematicians more effective may all deserve substantially more professional recognition than they have traditionally received.”
  • Custodian of epistemic self-government: An expert who helps a community independently evaluate, govern, and take responsibility for knowledge claims. “Experts are, in short, custodians of epistemic self-government.”
  • Epistemic redundancy: The presence of independent sources of expertise or judgment that allow claims to be checked, challenged, or overruled. “Expertise provides epistemic redundancy, allowing judgments made by machines to be interrogated, evaluated, and even overruled by experts.”
  • Epistemic self-government: The capacity of individuals or societies to evaluate and regulate knowledge claims independently. “A society whose members can receive correct answers to questions, but cannot independently criticize, interrogate, or understand the systems on which those answers depend, has surrendered a core part of its epistemic self-government.”
  • Epistemic status: The evidential significance or knowledge-related meaning assigned to a claim, object, or artifact. “One important and immediate consequence is that the epistemic meaning of a finished mathematical product has changed.”
  • Forensic surveillance: Intensive investigative monitoring intended to reconstruct or verify the origins and production of an artifact. “Such disclosure should be at a level appropriate to intellectual provenance, attribution, and reproducibility, rather than becoming an exercise in forensic surveillance.”
  • Formal verification: The use of formal logical systems to prove that a statement follows from specified assumptions and rules. “Formal verification answers a different question from understanding: whether a formal statement follows from specified assumptions under the rules of the system, rather than why an argument works, what its essential ideas are, or what the result means.”
  • Heuristic activity: Exploratory, experience-based reasoning used to discover or develop mathematical ideas without constituting a complete proof. “Modern philosophy of mathematical practice similarly emphasizes the centrality of mathematical agents, history, heuristic activity, and social epistemology.”
  • Human exceptionalism: The view that humans possess unique or superior status, abilities, or value relative to nonhuman entities. “Human exceptionalism, and nostalgia for a time when human control over society's destiny was more secure, are not the point.”
  • Independent mode: A mode of training in which students work without AI assistance when the goal is to develop or assess capacities that AI could perform. “In independent mode, students would work without reliance on artificial intelligence when the pedagogical aim is to develop or assess capacities that AI might otherwise perform for them.”
  • Institutional memory: Knowledge and practices retained by an institution across time and generations. “Among their distinguishing features are distributed expertise across many institutions and people; the preservation of subjects that are temporarily unfashionable; some insulation from immediate commercial utility; long time horizons that foster broader perspectives; institutional memory; and teaching as an integral part of learning.”
  • Intellectual agency: The capacity to independently select, evaluate, direct, interpret, and integrate intellectual work. “These capacities form the basis of intellectual agency.”
  • Intellectual culture: The shared practices, values, knowledge, and social environment through which a field’s expertise is developed and transmitted. “Intellectual culture is the environment in which such an apprenticeship takes place.”
  • Intellectual provenance: The origin and chain of intellectual contributions underlying a scholarly work or result. “Such disclosure should be at a level appropriate to intellectual provenance, attribution, and reproducibility, rather than becoming an exercise in forensic surveillance.”
  • Intergenerational: Occurring across successive generations or involving transmission between generations. “This axiom also connects explicitly to Husserl's later philosophy of mathematical practice, which views science as an intergenerational and intersubjective praxis.”
  • Intersubjective praxis: A socially shared form of practice constituted through interaction among conscious subjects. “This axiom also connects explicitly to Husserl's later philosophy of mathematical practice, which views science as an intergenerational and intersubjective praxis.”
  • Leiden Declaration: A set of proposed community principles concerning the responsible interaction between artificial intelligence and mathematics. “Several echo points in the Leiden Declaration on Artificial Intelligence and Mathematics \cite{Leiden}, and can be viewed as emerging community principles for interaction with artificial intelligence.”
  • Machine-checkable proof: A proof represented in a formal language that software can mechanically verify. “Machine-checkable proofs, including proofs formally verified in systems such as Lean, are not substitutes for understanding.”
  • Mathematical agency: The ability to make informed judgments, decisions, and interventions from within mathematical practice. “Intellectual agency is mathematical judgment exercised from within mathematics.”
  • Mathematical maturity: The developed ability to understand, evaluate, organize, and apply mathematical ideas beyond rote technical competence. “One could reasonably infer that its author was technically proficient; that they possessed the persistence, discipline, organization, and openness required to acquire substantial familiarity with a field; and that they had developed some degree of judgment and mathematical maturity, presumably under the guidance of an advisor.”
  • Mathematical ownership: Demonstrable understanding and responsibility for the mathematical content of a work, including AI-assisted work. “Glickenstein likewise argues that the dissertation defense can serve as a structural safeguard by assessing ``mathematical ownership'': a student who used AI in research or exposition should be able to explain what was used, what was checked or discarded, and why the final mathematics is correct.”
  • Natural generalization: An extension of a mathematical concept or result that preserves meaningful structural relationships rather than merely enlarging its scope artificially. “the maturity to distinguish natural from artificial generalizations.”
  • Normative: Relating to values, standards, or judgments about what ought to be done. “The point is normative rather than competitive: if human mathematical expertise, epistemic self-government, and a living human mathematical culture are goods worth preserving, then humans must continue to exercise these capacities even where machines can exercise them as well or better.”
  • Phenomenological: Relating to the philosophical study of conscious experience and how phenomena are experienced. “This process might be called digestion, and seems compatible with Husserl's phenomenological view of mathematical understanding.”
  • Praxis: Theory-informed, socially situated practical activity. “Modern philosophy of mathematical practice similarly emphasizes the centrality of mathematical agents, history, heuristic activity, and social epistemology.”
  • Premature offloading: Delegating a task to an external tool before developing the competence that performing the task is intended to build. “Premature offloading delegates a task whose performance is itself part of the process through which the relevant competence would have formed.”
  • Productive offloading: Delegating work to a tool in a way that removes unnecessary effort without undermining the competence needed for judgment. “Productive offloading delegates a task, cognitive or otherwise, to a machine in a way that removes unnecessary drudgery, expands capacity, or allows attention to be directed elsewhere without undermining the competence on which judgment depends.”
  • Proof abundance: A condition in which automated systems can produce very large numbers of mathematical proofs. “Tao has described one aspect of the current transition as a movement from proof scarcity'' towardproof abundance'' \cite{Tao}.”
  • Proof scarcity: A condition in which producing mathematical proofs is relatively difficult and limited by human labor. “Tao has described one aspect of the current transition as a movement from proof scarcity'' towardproof abundance'' \cite{Tao}.”
  • Proxy: An indirect measure used as evidence for a less directly observable property. “the dissertation can no longer, by itself, serve as a reliable proxy, much less sufficient evidence, for the formation of an independent mathematician.”
  • Sedimentation: The process by which meanings, practices, or achievements become deposited and preserved through historical development. “Connecting back to Husserl and the second axiom, through inscription and symbolic practice, whether human or machine, mathematical meanings become sedimented.”
  • Social epistemology: The study of how social interactions, institutions, and communities produce, validate, and transmit knowledge. “Modern philosophy of mathematical practice similarly emphasizes the centrality of mathematical agents, history, heuristic activity, and social epistemology.”
  • Stewardship: Responsible care, preservation, development, and transmission of a field’s intellectual resources and culture. “I propose that graduate training be understood as an apprenticeship toward stewardship of the discipline and the formation of intellectual agency.”
  • Synthetic writing: Writing that integrates and organizes multiple ideas, sources, or results into a coherent account. “These include research seminars, exposition, teaching, refereeing, collaborative work, AI-assisted exploration, structured mentoring of junior students, and synthetic writing.”
  • Tacit knowledge: Practical or contextual knowledge that is difficult to fully articulate or encode explicitly. “Some of the knowledge involved in this stewardship is tacit and situated rather than fully encoded in the literature.”
  • Technical proficiency: Sufficient internal mastery of a technical field to understand, evaluate, and work with its concepts and methods. “Technical proficiency supplies the internal reference point against which externally generated mathematics can be interpreted, criticized, reconstructed, and incorporated into one's own understanding.”
  • Technical virtuosity: Exceptional skill in executing difficult technical procedures or arguments. “Technical proficiency should not be confused here with technical virtuosity, nor with an ability to execute every difficult argument unaided.”

Open Problems

We're still in the process of identifying open problems mentioned in this paper. Please check back in a few minutes.

Tweets

Sign up for free to view the 2 tweets with 100 likes about this paper.