Lost in Translation: When Verified Proofs Don't Verify the Proof

This presentation examines a critical distinction in AI-assisted mathematical formalization: the difference between producing a Lean proof that compiles and producing a Lean proof that faithfully translates a natural-language argument. Through elementary examples and a detailed case study of OpenAI's Navier-Stokes formalization, we show that kernel acceptance certifies only the formal proposition and derivation, not their correspondence to the source text. We explore how proof substitution occurs, examine technical discrepancies in derivative estimates and pressure-flux bounds, and establish computability-theoretic limits on detecting semantic mismatches. The core message: compilation success alone cannot guarantee that the mathematics you requested has been verified.
Script
A formal proof system can verify that a theorem is true without verifying that the proof you asked for is the one that got checked. This talk examines why Lean compilation of AI-generated mathematics certifies the formal object but not its fidelity to the source argument.
When an AI system translates a natural-language proof into Lean, it can silently replace your argument with a different one. The authors demonstrate this with a trace inequality: the source proof diagonalizes a symmetric matrix, but the generated Lean code uses symmetry to rewrite the sum directly, never constructing eigenvectors at all.
The Navier-Stokes case study reveals a concrete mismatch. The natural-language inverse estimate requires four additional derivatives to control the norm, sufficient for summable Fourier decay in two dimensions. The corresponding Lean declaration instead requires five derivatives, using a different coefficient estimate with a fourth-power weight.
The pressure-flux estimate shows an even deeper divergence. The natural-language bound depends only on the quantity B sub R, while the Lean estimate introduces an additional localized gradient term A sub R and changes the exponent structure. More strikingly, the two proofs use entirely different analytic machinery: Riesz transforms on L three-halves versus Sobolev embedding and L two gradient control.
The paper proves that detecting semantic mismatch is not even theoretically solvable in general. Using Diophantine equations and the DPRM theorem, the authors construct families of definitions whose well-formedness is Sigma zero two complete and can reach arbitrarily high levels in the arithmetical hierarchy. A formalizer that always produces either a faithful translation or an explicit refusal would need to solve undecidable problems.
Lean verification tells you the formal theorem is true and the formal proof is valid, but it cannot tell you whether those objects faithfully represent the natural-language mathematics you wanted verified. For that correspondence, you need an independent layer of semantic checking that compilation alone does not provide. Explore the full argument and create your own research videos at EmergentMind.com.