Papers
Topics
Authors
Recent
Search
2000 character limit reached

Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs

Published 6 Oct 2026 in math.AP, cs.AI, and math.LO | (2610.08144v1)

Abstract: Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI's announced proof of blow-up of solutions to the Navier-Stokes equations. In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully. In particular, we highlight that the problem of resolving ambiguities in mathematical NL text, which is necessary in order to provide semantically faithful translation, is arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy/arithmetical hierarchy (the SCI =∞= \infty). Hence, informally, providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem (which has SCI =1= 1). To demonstrate the effect of this result we provide several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean `verifications'. These include OpenAI's announced Navier-Stokes proof. In particular, we show that the formalised Lean proof does not correspond to the NL proof of blow-up of solutions to the Navier-Stokes equations.

Summary

  • The paper highlights that successful Lean verification does not guarantee that formal proofs correspond to natural language proofs.
  • It identifies specific examples where Lean proofs differ from the intended arguments, particularly in OpenAI’s Navier–Stokes formalisation.
  • The authors demonstrate that semantic preservation requires matching not only the conclusion but the entire reasoning structure of the source, which cannot be universally ensured by compilation

Central thesis and scope

“Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs” (2610.08144) distinguishes two substantially different objectives that are often conflated under the term autoformalisation. The first is to obtain Lean code that proves a formally stated theorem. The second is to translate a natural-language (NL) mathematical text into Lean while preserving its semantic content, including definitions, hypotheses, intermediate claims, proof strategy, dependencies, and implicit existence conditions.

The paper’s central claim is that successful compilation addresses only the first objective. A Lean term that type-checks certifies the proposition encoded in Lean and the validity of the formal derivation supplied for that proposition. It does not certify that the encoded proposition is the one expressed in the source text, that the formal derivation follows the source argument, or that the source argument is correct. Consequently, a formal proof can coexist with an incorrect NL proof, or it can establish a different theorem by a different method.

The authors develop this distinction through elementary counterexamples, a detailed examination of two components of OpenAI’s announced Navier–Stokes formalisation, and computability-theoretic arguments concerning the complexity of semantic disambiguation. Their disclaimer is appropriately narrow: the paper does not establish that OpenAI’s NL proof is false. It establishes that the cited Lean formalisation is not, in the examples examined, a semantically faithful translation of the corresponding NL statements and arguments.

Two notions of autoformalisation

The weaker notion begins with a theorem that has already been translated into Lean and asks an AI system to produce any Lean proof accepted by the kernel. Under this criterion, a proof is successful if it compiles without sorry or additional axioms. The source proof functions only as guidance for proof search.

The stronger notion requires semantic preservation. A translation must preserve the mathematical meaning of the source text, not merely produce a proposition or proof with the same informal conclusion. This requires maintaining the source definitions, quantifier structure, side conditions, regularity assumptions, domains, norms, dependencies, and proof-relevant intermediate assertions. For a long mathematical text, the obligation is global: each formalised result must remain compatible with the meanings and dependencies established earlier.

This distinction has an immediate formal consequence. If the NL argument is invalid but the theorem statement is true, an AI system optimizing only for Lean acceptance can silently replace the invalid argument with a valid one. If the NL argument is valid, the system can still replace it with an unrelated formal proof. In either case, kernel verification provides no evidence that the argument requested for translation has been preserved.

The paper summarizes the resulting failure mode as follows:

Lean acceptance of a translation does not imply semantic preservation.

This is not a defect in Lean’s kernel. It is a mismatch between the kernel’s specification and the intended verification task. Lean verifies a formal object relative to its formal declarations; it does not compare that object with an external NL source.

Elementary examples of proof substitution

The first example uses the polynomial

p(x)=x3−x2−x+1.p(x)=x^3-x^2-x+1.

The source statement asserts nonnegativity on the relevant domain, but the accompanying NL proof assigns incorrect root multiplicities and therefore gives an incorrect factorization. The paper reports that the AI nevertheless generated a compiling Lean proof by discovering the correct algebraic structure. The result is a valid Lean proof of the theorem, but not a formalisation of the supplied NL proof.

Figure 1

Figure 1: An AI system translates an incorrect natural-language proof into a correct Lean proof by changing the mathematical argument.

This example isolates a particularly important asymmetry: a formaliser evaluated by compilation is incentivized to repair the source proof when repair is easier than faithful encoding. The resulting success may therefore be evidence that the theorem is provable, but it is not evidence that the original proof is sound.

The second example concerns a real symmetric matrix AA and the nonnegativity of tr⁡(A2)\operatorname{tr}(A^2). The NL proof diagonalizes AA in an orthonormal eigenbasis and then computes the trace as a sum of squared eigenvalues. The Lean proof instead retains the original basis and uses symmetry, aij=ajia_{ij}=a_{ji}, to rewrite

tr⁡(A2)=∑i,jaijaji=∑i,jaij2.\operatorname{tr}(A^2)=\sum_{i,j}a_{ij}a_{ji} =\sum_{i,j}a_{ij}^2.

Both proofs are mathematically valid, but the formal proof does not preserve the diagonalization argument.

Figure 2

Figure 2: A correct natural-language proof is translated into a different valid Lean argument that uses symmetry rather than diagonalization.

This example is not a correctness failure at the level of the theorem. It is a failure of proof correspondence. That distinction matters whenever the purpose of formalisation includes auditing a particular argument, checking the use of a delicate lemma, or preserving explanatory mathematical structure.

The Navier–Stokes case study

The paper applies this framework to OpenAI’s announced finite-time blow-up proof for the three-dimensional incompressible Navier–Stokes equations. The authors inspect the NL proof and the Lean repository at a specified commit, comparing source estimates and proof mechanisms with the declarations cited as their formal counterparts.

The first discrepancy concerns an inverse estimate in the source’s Lemma 8.6. The NL statement requires four additional derivatives: ∥N−1F∥Cym≤Cm∥F∥Cym+4,\|N^{-1}F\|_{C_y^m} \leq C_m\|F\|_{C_y^{m+4}}, whereas the corresponding Lean development requires five: ∥N−1F∥Cym≤Km∥F∥Cym+5.\|N^{-1}F\|_{C_y^m} \leq K_m\|F\|_{C_y^{m+5}}. Here N=vt⋅∇yN=v_t\cdot\nabla_y is a directional derivative on T2\mathbb{T}^2, and the inverse is defined on zero-mean functions.

Figure 3

Figure 3: The cited Lean inverse estimate requires one more input derivative than the corresponding natural-language estimate.

The difference is not a superficial notation change. The source argument uses Fourier decay sufficient to produce a summable AA0 factor in two dimensions after four additional derivatives. The Lean proof instead invokes a coefficient estimate involving a summable fourth-power weight, which leads to the AA1 requirement. The formal result is therefore weaker with respect to derivative loss. The authors also note that some formal statements are not directly comparable because they use different derivative types, including mixed derivatives rather than precisely the torus derivatives appearing in the NL norm.

The implication is specific: compilation of the Lean theorem does not establish the source’s sharper regularity estimate. At most, it establishes the formally stated estimate under stronger hypotheses. If the source proof depends quantitatively on the AA2 bound, the formalised result cannot be treated as its verification without an additional argument deriving the source estimate from the Lean development.

The second case concerns the pressure-flux estimate labelled (10.19) in the announced NL proof. The source gives a bound expressed purely through AA3, schematically

AA4

for almost every time. The corresponding Lean estimate has a different structure: AA5 and is asserted for every time in the relevant interval. The formal bound depends on the additional localized gradient quantity AA6, while the NL estimate does not.

Figure 4

Figure 4: The pressure-flux estimate and proof strategy in Lean differ from both the statement and proof in the natural-language argument.

The discrepancy arises from a substantive change in analysis. The NL proof uses boundedness of the Riesz transform on AA7, together with a decomposition involving AA8. The Lean proof instead uses Sobolev embedding to control an AA9 norm through an tr⁡(A2)\operatorname{tr}(A^2)0 gradient norm, applies tr⁡(A2)\operatorname{tr}(A^2)1 boundedness of the Riesz transforms, and uses a different Hölder decomposition involving tr⁡(A2)\operatorname{tr}(A^2)2 and tr⁡(A2)\operatorname{tr}(A^2)3. The derivative of the cutoff then produces the tr⁡(A2)\operatorname{tr}(A^2)4 contribution. The two arguments consequently establish different estimates.

The formal assertion that the estimate holds for every time, rather than almost every time, is also obtained through a different route: the Lean development establishes temporal continuity of the relevant integral. This may be a mathematically useful strengthening of a particular formal statement, but it does not repair the semantic mismatch. A stronger time-regularity conclusion combined with a different spatial estimate remains a different proposition and proof.

The paper therefore identifies two independent deviations in this case: the Lean code changes the bound and changes the method used to prove it. The source estimate cannot simply be recovered by observing that another source inequality relates tr⁡(A2)\operatorname{tr}(A^2)5 and tr⁡(A2)\operatorname{tr}(A^2)6; the authors explicitly state that the cited relation does not eliminate tr⁡(A2)\operatorname{tr}(A^2)7 from the Lean expression in a way that reproduces the NL estimate.

Why compilation is an insufficient acceptance criterion

The paper models compilation-driven autoformalisation as an iterative search process. Given a formal theorem and an NL proof, the system revises a candidate Lean proof until the kernel accepts it. The acceptance predicate checks only whether the candidate proves the formal theorem. It does not check whether the candidate preserves the source proof.

This creates three distinct outcomes:

  1. The NL proof is incorrect, while the generated Lean proof is correct.
  2. Both proofs are correct, but use different arguments.
  3. The NL text states a stronger claim than the Lean code proves.

All three outcomes satisfy the same compilation criterion. The third is especially consequential in analysis, where changes in derivative counts, integrability exponents, quantifiers over time, or dependence on auxiliary quantities can alter the applicability of subsequent estimates.

Back-translation from Lean to NL does not solve the problem automatically. One would still need to establish that the generated NL text faithfully represents the formal proof and that it corresponds to the original source argument. The semantic-preservation problem therefore reappears in the reverse direction.

Semantic ambiguity and arithmetical complexity

The paper’s broader theoretical claim is that semantic preservation is not merely difficult because mathematical language is informal. Some ambiguities depend on whether mathematical objects introduced by the source are well-defined, and that question can encode undecidable problems.

The construction uses an effective enumeration of integer-coefficient polynomials tr⁡(A2)\operatorname{tr}(A^2)8 and defines tr⁡(A2)\operatorname{tr}(A^2)9 as the least natural number AA0 for which

AA1

has no solution in natural numbers AA2. The source may then define AA3 and state an elementary identity involving AA4. The identity itself is trivial if AA5 exists. The semantic issue is whether the defining least number exists at all.

A formal system with default-valued minimization operators can silently assign a value when the intended set is empty. The paper gives a Lean-style construction using Nat.sInf, which returns AA6 on an empty set. The resulting code can prove the elementary ring identity while changing the source semantics: it has replaced a partial definition with a total one. Compilation therefore conceals, rather than resolves, the existence presupposition.

For AA7, the authors prove that the set of indices for which such an AA8 exists is AA9-complete, using the DPRM theorem and a nine-variable Diophantine representation of computably enumerable relations. Hence deciding whether the source definition is well-defined is not computable relative to the Halting problem. Generalized constructions yield aij=ajia_{ij}=a_{ji}0-complete problems for every finite aij=ajia_{ij}=a_{ji}1, with corresponding unbounded finite levels in the SCI hierarchy.

The paper summarizes this as an SCI value of infinity for the unrestricted problem of determining, across all such constructions, whether a source statement is semantically translatable or whether the formaliser must refrain from translating. The claim should be read in the paper’s precise sense: a universally trustworthy formaliser that always outputs either a faithful translation or an explicit refusal would have to solve ambiguity-resolution problems of arbitrarily high finite SCI. Thus, no fixed finite computational procedure can provide the proposed guarantee over the full class of source texts.

This result does not imply that every mathematical translation is computationally intractable, nor that useful semantic checks are impossible. It establishes a worst-case barrier for universal guarantees. Restricted mathematical languages, explicit well-formedness conditions, proof annotations, and human-supplied specifications may substantially reduce the problem instance, but those restrictions are external assumptions rather than consequences of Lean compilation.

Limitations and open questions

The paper’s empirical conclusions are based on selected examples rather than a systematic audit of the entire Navier–Stokes or Euler repositories. The authors explicitly state that they do not determine whether the OpenAI NL proof is correct, and they do not claim that the identified mismatches invalidate every formal theorem in the repository. Their conclusion is narrower and better supported: the cited Lean declarations do not faithfully formalize the particular NL claims and arguments examined.

The theoretical results concern universal semantic preservation over broad classes of mathematical texts. They do not specify a practical, restricted specification language in which faithful translation might be decidable or mechanically checkable. The paper also does not provide a complete protocol for certifying correspondence between an NL proof and a Lean proof. It leaves open how much correspondence can be established through proof annotations, structured intermediate representations, bidirectional translation, independently verified semantic parsers, or human-authored formal specifications.

A further open question is whether the source and formal developments could be connected by a richer notion of traceability: for example, a proof object augmented with machine-checkable links from each formal lemma to a precisely delimited NL assertion and its hypotheses. Such links would not remove the worst-case computability barrier, but the paper leaves open how effective they could be for restricted, practically important mathematical domains.

Conclusion

The paper establishes a precise separation between formal proof verification and semantic verification of autoformalisation. Lean can certify that a formal proposition follows from formal premises, but it cannot, without an independently checked correspondence layer, certify that those premises and that proof represent the intended NL mathematics.

The elementary examples demonstrate proof substitution, while the Navier–Stokes case study exhibits technically material discrepancies: an aij=ajia_{ij}=a_{ji}2 derivative requirement replacing an aij=ajia_{ij}=a_{ji}3 estimate, and a pressure-flux inequality with different quantities, exponents, time quantifiers, and proof mechanisms. The computability-theoretic analysis explains why such failures cannot be eliminated by a universal compilation-based procedure. The resulting methodological conclusion is direct: Lean-verified autoformalisation should be treated as verification of the formal artefact unless semantic faithfulness to the source has itself been established.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper asks an important question about using artificial intelligence to write mathematical proofs:

If an AI translates a proof into Lean code, and Lean accepts the code, does that prove that the original written proof was correct?

The authors argue that the answer is no.

Lean is a computer system that checks whether formal mathematical statements follow from earlier facts. If Lean accepts a piece of code, then the formal statement written in the code has been proved. However, the code might not faithfully represent the original proof written in ordinary language.

The paper uses examples involving simple algebra, matrices, and very advanced problems about the Navier–Stokes equations, which describe the movement of fluids such as water and air.

2. What questions do the authors investigate?

The paper mainly studies these questions:

  • Can an AI translate an ordinary mathematical proof into Lean code without changing its meaning?
  • Does a Lean proof always prove the same thing as the original written proof?
  • Can an AI accidentally fix, weaken, or completely change a proof while still producing code that Lean accepts?
  • How difficult is it to check that the translation has preserved the original meaning?
  • What happened when OpenAI translated its announced Navier–Stokes and Euler equation proofs into Lean?

The authors distinguish between two different tasks.

Translating a proof that Lean accepts

In the first task, an AI is given a theorem and a written proof. It tries to produce Lean code that proves the theorem. Success means only that the code compiles.

This is similar to asking a student to solve a problem and checking only whether the final answer is correct. The student may have used the wrong method, misunderstood the question, or even corrected an error in the original instructions.

Translating the meaning faithfully

The second task is much stricter. The AI must preserve:

  • the exact definitions,
  • the exact theorem,
  • the assumptions,
  • the logical steps,
  • and the overall mathematical argument.

This is like translating a book into another language while making sure that every idea, detail, and connection remains unchanged. According to the authors, this is much harder than simply producing code that passes Lean's checker.

3. How did the authors investigate the problem?

The authors use several approaches.

Simple mathematical examples

They first examine small examples where it is easy to see what went wrong.

For instance, they describe a polynomial proof with an incorrect explanation. The original written proof gives the wrong roots and an incorrect factorisation. However, the AI produces Lean code that proves the correct result by silently using the correct roots.

So Lean checks a correct proof, but not the incorrect proof that the AI was supposed to translate.

They also give a matrix example. The original proof uses a special change of coordinates called diagonalisation. The AI's Lean proof proves the same result using a different idea based on the symmetry of the matrix.

Both proofs are valid, but the Lean proof is not a faithful translation of the original proof.

Comparing ordinary text with Lean code

The authors then examine the natural-language proofs and the associated Lean files from OpenAI's announced Navier–Stokes and Euler projects.

They compare:

  • what the written papers claim,
  • what the Lean theorems actually assume,
  • what the Lean theorems actually prove,
  • and which mathematical methods are used.

This is similar to comparing an original set of building instructions with a computer-generated construction plan. The authors check whether the final plan builds the same structure using the same important requirements.

Studying the difficulty of faithful translation

The paper also uses ideas from logic and computer science, including the Solvability Complexity Index hierarchy and the Halting Problem.

These are ways of measuring how difficult certain computational questions are. The Halting Problem asks whether a computer program will eventually stop or continue running forever. It is known that no general algorithm can always answer this correctly.

The authors argue that deciding whether an AI translation has preserved the exact meaning of a mathematical text can be even more difficult in a formal sense. Their point is not that every translation is impossible, but that there is no simple universal checking method that can always guarantee faithful translation.

4. What did the authors find?

Lean can verify a different proof

The main finding is that Lean may verify a proof that is:

  • based on a different argument,
  • about a weaker statement,
  • based on different assumptions,
  • or unrelated to the original written reasoning.

A proof being accepted by Lean therefore gives confidence only in the formal Lean statement, not automatically in the natural-language proof from which it was supposedly created.

Example: a wrong written proof becomes a correct Lean proof

In the polynomial example, the written argument is wrong, but the AI generates a correct Lean proof. This makes the code look trustworthy even though it did not translate the original reasoning.

The AI effectively changed the proof instead of translating it.

Example: a correct proof becomes a different proof

In the matrix example, the written proof is correct. However, the Lean code leaves out an important step and uses another method.

This shows that even a correct translation can still be unfaithful if the goal was to preserve the original argument.

The Navier–Stokes formalisation uses different statements

The paper focuses especially on an estimate in OpenAI's announced Navier–Stokes proof.

In the written paper, one estimate requires control of derivatives up to roughly order m+4m+4. In the Lean code examined by the authors, the corresponding result requires derivatives up to roughly order m+5m+5.

In everyday language, this means that the Lean theorem asks for more information about the function than the written proof does. It may still be a correct theorem, but it is a different and generally weaker result.

An analogy would be:

  • The original instructions claim that a bridge can hold a certain weight using four support beams.
  • The formal version proves that it can hold the weight if five support beams are used.

The second claim may be true, but it does not prove the first claim.

Another Navier–Stokes argument is substantially different

The paper also examines a pressure-related estimate. The written proof and the Lean proof use different formulas, different assumptions, and different mathematical tools.

For example, the Lean proof uses techniques involving:

  • Sobolev embedding, which connects different ways of measuring the size or smoothness of a function;
  • Hölder's inequality, a rule for estimating products;
  • Riesz transforms, special operations used in the analysis of functions and fluid equations.

These methods are not necessarily wrong. The problem is that the Lean proof does not appear to be a faithful formal version of the written proof.

The resulting estimate also contains extra quantities that do not appear in the natural-language version. Thus, the two proofs may establish different results.

Lean compilation does not check the meaning of names and comments

Lean checks the formal definitions and logical steps written in its code. It does not know whether those definitions correctly match the explanations in a paper.

For example, if a paper says “prove statement A,” but the Lean code formalises statement B, Lean can correctly verify B without noticing the mismatch.

5. Why are these findings important?

Formal proof systems such as Lean are extremely useful. They can catch many ordinary mathematical mistakes, such as missing assumptions or invalid algebraic steps. But they cannot automatically guarantee that an AI translated a human-written proof correctly.

The paper therefore warns against treating a compiled Lean file as automatic proof that the original paper is correct.

To trust the complete result, mathematicians still need to check:

  1. whether the Lean definitions match the original definitions;
  2. whether the formal theorem says the same thing as the written theorem;
  3. whether the formal assumptions match the assumptions in the paper;
  4. whether the Lean proof follows the same important ideas;
  5. and whether the translation has accidentally weakened or changed the result.

Conclusion: What could this mean for the future?

AI and formal proof systems could become powerful tools for mathematics. They may help researchers find errors, explain proofs, and check complicated calculations.

However, this paper shows that there are two separate questions:

  • Is the Lean code correct?
  • Does the Lean code faithfully represent the original mathematical argument?

Lean can answer the first question, but not automatically the second.

The likely best use of AI formalisation is therefore as an aid to mathematicians, not as a replacement for careful reading and peer review. Human experts must still compare the original proof with the formal code.

The paper's main message is simple:

A computer-checked proof is only as trustworthy as the connection between the computer's formal statement and the original mathematics.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

The paper leaves the following issues unresolved:

  • No general, implementable criterion for semantic faithfulness is provided. The paper argues that faithful autoformalisation is computationally intractable in general, but does not specify practical sufficient conditions under which a particular NL-to-Lean translation can be trusted.
  • The SCI/arithmetical-hierarchy result is not fully developed in the presented text. The exact formal decision problem, encoding of mathematical texts and meanings, reduction establishing arbitrary SCI height, and assumptions needed for the claim that semantic disambiguation has SCI =∞=\infty are not stated in sufficient detail to independently assess or reproduce the result.
  • The relationship between semantic-faithfulness checking and undecidability is not sharply delimited. It remains unclear which restricted classes of mathematical language, proof styles, domains, or formal systems might admit decidable or semi-decidable faithfulness checks.
  • No quantitative notion of semantic similarity or proof correspondence is defined. The paper identifies translations that are weaker, stronger, or mathematically different, but does not formalise how much divergence is permissible or how correspondence between individual NL steps and Lean terms should be measured.
  • There is no validated automated method for detecting the reported mistranslations. The examples are found through detailed manual comparison, but the paper does not present an algorithm, benchmark, or tool that can systematically identify omitted arguments, altered hypotheses, changed derivative orders, or substituted proof strategies.
  • The empirical evidence is narrow. The practical analysis focuses mainly on selected examples from OpenAI’s Navier–Stokes formalisation, with only brief mention of the Euler formalisation and elementary examples; it does not establish how frequent or representative such mistranslations are across models, mathematical fields, theorem provers, or prompting strategies.
  • The paper does not evaluate competing autoformalisation systems. It remains unknown whether the observed failures are specific to the systems and repositories examined or reflect a broader limitation shared by current LLMs and formalisation pipelines.
  • The effects of model version, prompting, sampling, and post-processing are not studied. The paper does not report whether mistranslations persist across repeated generations, different prompts, model temperatures, model updates, or human-assisted correction workflows.
  • The analysis does not provide a complete audit of the OpenAI Navier–Stokes or Euler repositories. The two detailed discrepancies demonstrate non-faithfulness, but the paper does not determine the full number, severity, or mathematical consequences of all mismatches in the formalised developments.
  • The downstream impact of the identified mismatches on the main theorems is not established. In particular, the paper shows that intermediate NL claims and Lean statements differ, but does not fully determine whether the Lean development still proves a meaningful substitute theorem, whether the NL proof can be repaired, or whether either development establishes the advertised blow-up result.
  • The comparison of the m+4m+4 and m+5m+5 derivative estimates is not completed mathematically. The paper identifies a weaker Lean estimate but does not prove whether the NL estimate is valid under all stated assumptions, whether the extra derivative is genuinely necessary for the Lean proof, or whether the Lean formalisation could be strengthened to recover the NL bound.
  • The pressure-flux comparison does not resolve equivalence or implication between the two estimates. Because the Lean bound depends on ARA_R while the NL bound is expressed using BRB_R, the paper leaves open whether additional hypotheses or interpolation inequalities could make the bounds comparable.
  • The correctness of the NL Navier–Stokes arguments is deliberately left unassessed. The paper establishes mistranslation rather than proving that the original NL estimates or proof steps are false, valid, or sufficient for the final theorem.
  • The treatment of informal mathematical conventions is incomplete. Issues such as suppressed quantifiers, “almost everywhere” versus pointwise claims, dependence of constants, domain conventions, and implicit regularity assumptions are discussed selectively rather than systematically catalogued.
  • No method is given for preserving proof provenance. The formalisation pipeline does not appear to record which Lean hypotheses, lemmas, and tactics correspond to each NL sentence or mathematical inference, leaving unresolved how auditors could trace a compiled proof back to the source argument.
  • The paper does not distinguish semantic faithfulness from mathematical equivalence rigorously enough for automated use. Two proofs may use different but equivalent arguments, yet the paper does not define when a changed proof strategy should count as an acceptable translation rather than a mistranslation.
  • The role of human verification remains unspecified. The paper recommends peer review and scrutiny but does not determine what level of human expertise, time, or documentation is required to validate a large autoformalised text.
  • No benchmark suite for faithful autoformalisation is proposed. Future work lacks a standard collection of NL statements and proofs annotated with intended meanings, acceptable formal equivalents, known ambiguities, and adversarial mistranslations.
  • The generality of the conclusions across formal systems is untested. The discussion centres on Lean, so it remains open whether analogous failures occur in Isabelle, Coq, Agda, Metamath, or systems using different type theories and proof representations.
  • The paper does not investigate safeguards that could reduce mistranslation risk. Possible interventions—interactive clarification, proof-step alignment, bidirectional translation, independent formalisation, theorem-statement locking, or semantic model checking—are mentioned neither experimentally nor theoretically.
  • The security and reliability implications of accepting compilable but semantically altered proofs are not analysed. The paper does not examine how such failures might affect mathematical databases, automated theorem corpora, scientific software, or AI systems trained on formally verified but source-inconsistent proofs.

Practical Applications

Immediate Applications

  • Research software: separate compilation from semantic verification. Mathematics and AI tooling teams can immediately treat Lean compilation as evidence only that the formalized proposition and formal proof are valid—not that they faithfully represent the source natural-language (NL) argument. A practical workflow is to maintain two explicit validation stages:

    1. verify that the Lean code compiles without sorry or unintended axioms;
    2. independently check that definitions, theorem statements, intermediate lemmas, hypotheses, and proof strategy correspond to the NL source. This applies directly to theorem-proving systems, proof assistants, and AI-generated mathematical repositories. It depends on reviewers having access to both source prose and formal code.
  • Mandatory source–formalization audits for AI-generated proofs.

    • the exact strength of the NL theorem with the Lean theorem;
    • derivative orders, quantifiers, domains, regularity assumptions, and exceptional sets;
    • the variables and signs used in corresponding expressions;
    • the mathematical argument, rather than only the final conclusion.
    • The paper’s Navier–Stokes examples show why this is actionable: a formal estimate requiring m+5m+5 derivatives can be weaker than an NL claim requiring m+4m+4, and a bound involving additional quantities such as ARA_R may not establish the stated NL bound.
  • Use formal proof assistants as theorem checkers, not translation validators. Researchers can safely deploy Lean, Coq, Isabelle, or similar systems to check a previously specified formal theorem. They should not describe successful compilation as validation of an AI-generated translation unless semantic correspondence has also been established. This distinction is immediately relevant to mathematical publishing, grant-funded software, and institutional reproducibility standards.
  • Repository-level provenance and dependency tracking.
    • commit hashes and versioned source documents;
    • links between each NL statement and its formal declaration;
    • a record of definitions and imported lemmas;
    • explicit declarations of added assumptions and weakened conclusions;
    • automated checks that no sorry, hidden axiom, or unintended theorem replacement is used.
    • Such metadata would make discrepancies like the paper’s m+4m+4 versus m+5m+5 derivative mismatch easier to detect.
  • Adversarial benchmark datasets for autoformalization systems.
    • incorrect NL proofs of correct statements;
    • correct proofs that can be replaced by different valid arguments;
    • statements whose formal versions are weaker or stronger than the source;
    • ambiguous notation, omitted quantifiers, domain changes, and altered regularity assumptions.
    • Systems should be evaluated not only on compilation or theorem-proving success, but also on semantic alignment, proof-step correspondence, and preservation of theorem strength.
  • Human-in-the-loop review interfaces for mathematical code.
    • omitted changes of basis;
    • replacement of an L3/2L^{3/2} argument with an L2L^2-based Sobolev argument;
    • altered “almost everywhere” versus “for all” quantifiers;
    • extra hypotheses or different norm definitions.
    • The method is deployable now, although reliable semantic-difference detection remains partly manual.
  • Improved academic peer-review checklists.
    • which portions were generated or translated by AI;
    • whether the NL proof was independently reviewed;
    • whether the formal theorem is equivalent to the published claim;
    • all additional assumptions and weakened intermediate results.
    • This follows directly from the paper’s conclusion that formal code does not eliminate the need for conventional mathematical peer review.
  • Educational training in formal-methods literacy.
    • a correct formal proof may prove a different argument;
    • a compiler verifies code against formal definitions, not authorial intent;
    • theorem equivalence and proof correspondence are separate questions.
    • This is immediately applicable in courses on proof assistants, mathematical logic, software verification, and responsible AI.
  • Safer interpretation of AI-generated scientific claims. Scientists, journalists, funding bodies, and policy organizations can avoid presenting “formally verified” AI mathematics as independently confirmed science. For claims involving fluid dynamics, physics, or engineering, formal verification should be reported with its precise scope: which formal theorem was checked, which assumptions were encoded, and whether the original exposition was semantically matched.
  • Formalization of localized analytic estimates in applied mathematics. The Lean developments discussed in the paper may still be useful as reusable libraries for Fourier analysis, Riesz transforms, Sobolev estimates, torus operators, pressure estimates, and Navier–Stokes analysis. Researchers can use these components in new work, provided each theorem’s exact hypotheses and conclusion are checked. The immediate value is in machine-checked sublemmas and reusable proof infrastructure, not automatic confirmation of the surrounding NL proof.

Long-Term Applications

  • Semantics-aware autoformalization systems. A future generation of systems could jointly produce:

    1. a formal theorem and proof;
    2. a structured semantic representation of the NL source;
    3. a proof-alignment certificate showing how each source definition, hypothesis, and inference maps to formal code. Such a system would need to establish equivalence or conservativity between source and formal statements, not merely generate compilable code. The paper emphasizes that unrestricted faithful interpretation of mathematical language can be arbitrarily high in the SCI/arithmetic hierarchy, so a universally reliable solution is not expected. Practical systems will therefore require restricted domains, controlled languages, or human certification.
  • Controlled mathematical languages for high-assurance translation.

    • quantifier scope;
    • domains and codomains;
    • norm conventions;
    • regularity and integrability assumptions;
    • exceptional-set qualifiers;
    • dependency relationships between results.
    • These languages could compile into Lean or another proof assistant and substantially reduce ambiguity. Adoption would depend on balancing formal precision with usability for mathematicians.
  • Equivalence-checking tools for theorem statements.
    • missing hypotheses;
    • extra derivatives;
    • changed norms or exponents;
    • altered time quantifiers;
    • dependence on previously absent quantities;
    • differences between almost-everywhere and pointwise claims.
    • This would be particularly valuable in PDEs, functional analysis, probability, and optimization, where small changes in regularity assumptions can materially alter a result. It requires advances in mathematical language understanding and formal statement normalization.
  • Proof-structure and proof-intent verification. Future systems could test whether a formal proof preserves the intended strategy of an NL proof—for example, whether a claimed diagonalization argument actually uses a change of basis, or whether a pressure estimate uses the stated Riesz-transform bound rather than a different Sobolev argument. This is more demanding than checking theorem equivalence because multiple valid proofs may exist. Feasibility depends on formalizing proof plans, intermediate goals, and acceptable strategy variations.
  • Certified AI-assisted mathematical publishing pipelines.
    • the readable paper;
    • machine-checkable formal statements;
    • proof code;
    • semantic alignment certificates;
    • a dependency graph;
    • reproducible build instructions;
    • human review records.
    • This could become a standard for major mathematical claims, especially in fields where computational or AI-generated arguments are increasingly used. It requires community standards, stable proof-assistant ecosystems, and publisher support.
  • Regulatory and institutional standards for high-impact AI mathematics.
    • formal compilation: the encoded theorem has a checked proof;
    • formal statement validation: the encoded theorem matches the intended claim;
    • proof correspondence: the formal proof reflects the published argument;
    • independent mathematical review: experts assess novelty and correctness.
    • This tiered framework would prevent the conflation identified by the paper while preserving the benefits of formal verification.
  • Automated discrepancy mining across large formal repositories.
    • formal theorems with systematically stronger hypotheses than the prose;
    • unused definitions or lemmas from the NL proof;
    • formal declarations that are never used in the final theorem;
    • unexplained changes in constants, exponents, derivative counts, or domains;
    • proof scripts that bypass the claimed intermediate results.
    • Such systems could prioritize files for expert review, making large-scale auditing more practical.
  • Domain-specific verification for engineering and scientific computation.
    • fluid dynamics: verified discretization assumptions and PDE estimates;
    • robotics and control: formal correspondence between natural-language safety claims and controller specifications;
    • finance: verification that risk models encode the stated regulatory or economic assumptions;
    • energy systems: machine-checked optimization constraints and stability claims;
    • healthcare: verification that clinical decision rules match published protocols.
    • These applications require domain-specific ontologies, validated models, and careful treatment of uncertainty; compilation alone would be insufficient evidence of real-world safety.
  • Personal and everyday use of trustworthy mathematical assistants. In the longer term, students, engineers, and the public could use assistants that explain a proof in natural language, formalize it, and explicitly report any semantic uncertainty or altered assumptions. Such assistants could support homework, spreadsheet reasoning, technical documentation, and decision support. Their feasibility depends on transparent uncertainty reporting, restricted deployment domains, and safeguards against presenting a formally valid but semantically unrelated argument as an answer.
  • Fundamental research on the limits of AI reasoning and translation.
    • syntactic proof checking;
    • theorem proving;
    • semantic interpretation;
    • proof-intent preservation.
    • This may lead to formal impossibility results, restricted completeness theorems, and better design principles for AI systems that distinguish what can be mechanically verified from what must remain subject to human interpretation and mathematical judgment.

Glossary

  • Arithmetical hierarchy: A classification of decision problems according to the complexity of the quantifiers required to define or solve them. “the Solvability Complexity Index (SCI) hierarchy/arithmetical hierarchy”
  • Autoformalisation: The automated translation of mathematical statements or proofs from natural language into a formal system. “Autoformalisation is increasingly used to verify mathematical texts”
  • Commutator: An operator measuring the failure of two operators to commute under composition. “where [⋅,⋅][\cdot,\cdot] is the standard commutator”
  • ContDiff: A formal-analysis predicate indicating that a function has continuous derivatives up to a specified order. “(hf : ContDiff ℝ ∞ f)”
  • Distribution: A generalized function defined through its action on test functions, allowing derivatives to be interpreted weakly. “The proof also defines a distribution π∗=∑i,jRiRjgij\pi_* = \sum_{i, j} R_i R_j g_{ij}”
  • Eigenbasis: A basis consisting of eigenvectors of a linear operator or matrix. “Diagonalise AA in an orthonormal eigenbasis.”
  • Fourier multiplier: An operator defined by multiplying the Fourier transform of a function by a specified symbol. “Let RiR_i be the Riesz transform with Fourier multiplier operator iξi/∥ξ∥2i\xi_i/\|\xi\|_2.”
  • Fourier series: A representation of a periodic function as a sum of sinusoidal or complex exponential modes. “Four more derivatives of FF than the requested output leave a summable (1+∣k∣)−3(1+|k|)^{-3} bound on the Fourier series in two dimensions.”
  • Haar mean: The integral or average associated with Haar measure on a locally compact group, often used for periodic domains. “with zero normalized Haar mean”
  • Halting problem: The undecidable problem of determining whether an arbitrary program eventually terminates. “harder than any computational problem including the Halting problem”
  • Hypodissipative: Characterizing an evolution equation whose dissipative term is weaker than the standard Laplacian dissipation. “CordobaMartinezZoroaZheng2026Hypodissipative”
  • Incompressible: Describing a velocity field whose divergence is zero, corresponding to conservation of volume. “The three-dimensional incompressible Navier-Stokes regularity problem”
  • Interpolation inequality: An inequality that bounds a norm between two endpoint norms by combining them with suitable exponents. “an interpolation inequality \href{https://github.com/openai/NavierStokesAndEuler/blob/f9e8bc5b38b6e212696e8a30e3e91517af887bbd/NavierStokes/R3/WeightedInterpolation.lean#L150-L161}{cutoff\_interpolation\_twelve\_fifths}”
  • Lean: A proof assistant and dependently typed programming language used to encode and mechanically verify formal mathematics. “the argument expressed in the formal language can easily be mechanically verified”
  • Leray solution: A weak solution of the Navier–Stokes equations satisfying appropriate energy bounds, named after Jean Leray. “dating back to Leray's fundamental paper from 1934”
  • Localized norm: A norm computed after restricting or weighting a function to a spatial region. “so that ARA_R is a localized L2L^2 gradient norm and BRB_R a localized L6L^6 velocity norm.”
  • Mean-zero solution: A solution constrained to have zero average, often to eliminate an additive-constant ambiguity. “The zero-mean requirement assures this solution is uniquely specified”
  • Multi-index: A tuple of nonnegative integers used to represent mixed partial derivatives. “where ww is a multi-index.”
  • Navier–Stokes equations: Nonlinear partial differential equations governing viscous fluid motion. “the Navier-Stokes equations”
  • Orthonormal eigenbasis: An eigenvector basis whose vectors are mutually orthogonal and have unit norm. “Diagonalise AA in an orthonormal eigenbasis.”
  • Pressure flux: The contribution of pressure to the transport of fluid momentum through a region or boundary. “The NL paper's pressure-flux bound and its proof are both mistranslated”
  • Riesz transform: A singular integral operator whose Fourier multiplier is proportional to a coordinate divided by frequency magnitude. “The Lean proof does not rely on the boundedness of the Riesz transform”
  • Schwartz function: A smooth function whose derivatives decay faster than any polynomial at infinity. “the boundedness of the Riesz operators from Schwartz functions in L2L^2 to functions in L2L^2”
  • Semantically faithful translation: A formal translation that preserves the mathematical meaning, dependencies, and inferential structure of the source text. “Faithful semantic translation of a mathematical text.”
  • Sobolev embedding: A theorem relating Sobolev-space regularity to integrability or continuity in another function space. “The Lean proof does not rely on the boundedness of the Riesz transform as a map from L3/2→L3/2L^{3/2} \to L^{3/2}; it instead invokes Sobolev embedding”
  • Solvability Complexity Index (SCI): A hierarchy classifying the number and type of limiting computational procedures needed to solve a mathematical problem. “the problem of resolving ambiguities in mathematical NL text ... is arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy”
  • Summable series: An infinite series whose terms have a finite total sum. “leave a summable (1+∣k∣)−3(1+|k|)^{-3} bound”
  • Symmetric matrix: A matrix equal to its transpose, satisfying aij=ajia_{ij}=a_{ji}. “Let A=(aij)A = (a_{ij}) be a real symmetric n×nn\times n matrix.”
  • Tensor: A multidimensional mathematical object whose components transform according to specified rules under changes of coordinates. “where $g = (g_{ij})_{i,j \in \{1,2,3\}$”
  • Torus: A space obtained by identifying opposite sides of a rectangle, commonly represented as Rn/Zn\mathbb{R}^n/\mathbb{Z}^n. “with T2:=R2/Z2\mathbb{T}^2:= \mathbb{R}^2/\mathbb{Z}^2”
  • Weak solution: A function satisfying a differential equation in an integrated or distributional sense rather than through pointwise derivatives. “In fact, the purpose of the lemma is to establish that w=0w = 0”
  • Zero Fourier coefficient: The constant Fourier mode of a periodic function. “uniqueness after fixing the zero Fourier coefficient”
  • Zero-mean condition: A constraint requiring the integral or average of a function over its domain to vanish. “with zero normalized Haar mean”

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 6 tweets with 172 likes about this paper.

HackerNews

  1. Navier–Stokes Lost in Translation (239 points, 150 comments)