---
title: Proof-State Snapshotting in Lean 4
url: https://www.emergentmind.com/topics/proof-state-snapshotting
type: topic
---

# Proof-State Snapshotting in Lean 4

Proof-state snapshotting is a technique for automated theorem proving in Lean 4 that captures an already elaborated proof state at a `sorry` position and reuses it across multiple tactic-search branches instead of reconstructing that state independently for each branch. In the formulation reported by Shen and Shi, the relevant proof state at source position $p$ is the triple $\mathrm{ProofState}(p) \coloneqq (\mathrm{Environment}_p,\ \mathrm{MetavarContext}_p,\ \mathrm{LocalContext}_p)$, and Lean’s implementation bundles these components into a `Snapshot` object. The method was introduced to address the mismatch between portfolio-style tactic search and an execution model in which each branch re-runs import loading and theorem-body elaboration, an arrangement that makes branching disproportionately expensive in Lean 4 with Mathlib [2605.25556].

## 1. Definition and internal representation

In Lean 4 with Mathlib loaded, a proof state at position `p` is represented by three internal data structures: the `Environment`, the `MetavarContext`, and the `LocalContext` [2605.25556]. The `Environment` contains the global environment of all loaded constants, declarations, notation, and type-class instances; the `MetavarContext` contains the current open metavariables and their assignment table; and the `LocalContext` contains the local hypotheses and bound variables visible at the relevant source position.

The paper states that these three components are bundled into a `Snapshot` object with the simplified type signature:

```lean
structure Snapshot :=
  (env    : Environment)
  (mctx   : MetavarContext)
  (lctx   : LocalContext)
```

A snapshot at position $p$ is then written as `snap_p : Snapshot = (env_p, mctx_p, lctx_p)` [2605.25556]. The `Environment` is described as approximately 2–4 GB once Mathlib is in core, whereas the `MetavarContext` is approximately KB per hole. This size asymmetry is central to the design: the large global environment is immutable and can be shared by reference, while the much smaller per-branch state can be cloned cheaply.

The same source also formalizes an implementation-level reference object:
$\mathrm{SnapshotRef} \hookrightarrow \{ \mathrm{envRef} : \ell_e,\ \mathrm{mctxRef} : \ell_m,\ \mathrm{lctxRef} : \ell_l \}$,
where $\ell_e$ points into the global read-only environment store and $\ell_m,\ell_l$ point into the proof’s mutable context store. This representation allows a JSON-RPC interface to return a stable `snapshotId : String` without transmitting the full proof state through JSON. A plausible implication is that the technique is best understood not as proof serialization, but as server-side retention and controlled reuse of elaborated state.

## 2. Baseline reconstruction and the source of overhead

The baseline targeted by proof-state snapshotting is the “Level 0” tactic-search workflow in which every branch rebuilds a fresh proof state [2605.25556]. The paper decomposes that reconstruction into two dominant costs.

First, import loading deserializes all Mathlib `.olean` files into an empty `Environment`, with cost $T_{\mathrm{load}} \approx 60\ \mathrm{s}$ per branch. Second, theorem-body elaboration re-type-checks the theorem header and all prior steps up to the first `sorry`, with cost $T_{\mathrm{body}} \in [18\ \mathrm{s}, 735\ \mathrm{s}]$ per branch, growing with proof complexity. Together these account for more than $99\%$ of per-branch wall time in a standard single-machine setting [2605.25556].

The paper gives concrete endpoint estimates. On the simplest benchmark, `mathd_algebra_209`, the minimum theorem-body elaboration cost is approximated as $\min(T_{\mathrm{body}}) \approx 78\ \mathrm{s} - 60\ \mathrm{s} \approx 18\ \mathrm{s}$. On the hardest cited case, `mathd_numbertheory_328` with 28 branches, theorem-body elaboration is estimated as
$T_{\mathrm{body}} \approx (11\,127\ \mathrm{s} - 28 \cdot 60\ \mathrm{s})/(\lceil 28/2 \rceil) \approx 735\ \mathrm{s}$.
Since tactic execution itself is only a few ms to a few hundred ms, the fallback time is summarized as
$T_{\mathrm{fallback}} \approx T_{\mathrm{load}} + T_{\mathrm{body}} + T_{\mathrm{tactic}} \approx 75\text{–}795\ \mathrm{s}$ per branch [2605.25556].

The paper’s per-branch breakdown illustrates the imbalance:

| Problem | Branches | Tactic CPU | Native/branch | Overhead% |
|---|---:|---:|---:|---:|
| `mathd_numbertheory_3` | 21 | 44 ms | 5.5 s | 99.9 % |
| `algebra_sqineq_at2malt1` | 14 | 50 ms | 5.7 s | 99.9 % |

These figures support the paper’s central diagnosis: the limiting factor is not tactic evaluation but repeated reconstruction of a largely identical proof state. This suggests that improvements to branching throughput should target state reuse rather than tactic-level micro-optimization.

## 3. Snapshot capture, branching, and execution model

To eliminate repeated reconstruction, the paper extends the Lean 4 Language Server with three JSON-RPC methods registered on server startup [2605.25556]. These endpoints are:

1. `$/lean/dspSnapshotPing`, which checks for snapshot support and replies `{ "ok": true }` if the patched server is running.
2. `$/lean/dspSnapshotCapture`, which takes a file `uri` and a `position` corresponding to a `sorry`, blocks until the internal elaborator reaches that position, records a handle to the corresponding `Snapshot`, and returns a stable `snapshotId : String`.
3. `$/lean/dspSnapshotBranch`, which takes a `snapshotId` and an array of tactic configurations, fetches the recorded snapshot, and spawns a parallel task for each tactic.

For each branch, the system clones only the `MetavarContext`, substitutes the tactic in place of `sorry`, runs the tactic, and returns a record of the form `{ ok: Bool, error: Option String, cpuSeconds: Float }` [2605.25556]. The `Environment` is shared by reference because it is immutable and much larger than the mutable proof-state components. The paper therefore characterizes the forking cost as $O(\mathrm{KB})$ and “effectively free even for hundreds of branches.”

The server-side procedure is given in LaTeX-style pseudocode:

$$
\begin{algorithmic}[1]
\Function{dspSnapshotBranch}{snapshotId, configs}
  \State snap \leftarrow lookupSnapshot(snapshotId)
  \ForAll{cfg in configs}
    \State \_ \leftarrow IO.asTask \Comment{spawn in parallel}
    \State mctx' \leftarrow snap.mctx.clone()
    \State lctx' \leftarrow snap.lctx.clone()
    \State (ok, err, cpu) \leftarrow runTactic(snap.env, mctx', lctx', cfg.tactic)
    \Return \{ ok := ok, error := err, cpuSeconds := cpu \}
  \EndFor
  \State results \leftarrow awaitAllTasks()
  \Return results
\EndFunction
\end{algorithmic}
$$

Each branch runs in parallel on a worker thread using Lean’s work-stealing scheduler, and wall time is described as dominated by the slowest single branch plus approximately 10 ms for task dispatch [2605.25556]. In effect, snapshotting changes the unit of reuse from files and imports to the fully elaborated proof state at the hole boundary.

## 4. Experimental protocol and benchmark design

The evaluation was performed on a MacBook Pro with Apple M-series, 8 GB RAM, and macOS, using Lean 4.25.1 patched via `elan` [2605.25556]. The data set comprised 48 miniF2F-v2 problems: 45 hand-crafted “Prove-phase” sketches with 1–5 `sorry` holes each, and 3 full end-to-end Draft-Sketch-Prove runs on selected theorems.

For each problem, the paper compares two paths. The native snapshot path uses the sequence `dspSnapshotPing` $\rightarrow$ `dspSnapshotCapture` $\rightarrow$ `dspSnapshotBranch`. The fallback path performs a full `lake build` per branch with no snapshots [2605.25556]. Fallback concurrency is bounded by RAM: with 8 GB, the experiments used approximately $W \approx 2$ parallel `lake` builds for the hand-crafted benchmarks and $W = 1$ for the full DSP runs.

Both paths use the same tactic portfolio of seven tactics—`aesop`, `norm_num`, `omega`, `ring`, `linarith`, `decide`, and `simp`—on the same proof sketches, so the measurement isolates infrastructure overhead rather than tactic quality [2605.25556]. The paper denotes the number of holes by $H$ and sets the number of branches to $B = H \times 7$.

This design matters methodologically because it separates speed from proving power. The system does not compare different search heuristics or different theorem-proving capabilities; it measures the cost of trying the same set of tactics under two different execution models.

## 5. Empirical results and scaling behavior

Across 48 miniF2F-v2 problems—45 Prove-phase benchmarks and 3 full end-to-end runs—the reported wall-time speedup over the standard fallback ranges from $5.6\times$ to $50\times$, with average $14\times$ and median $9.7\times$ [2605.25556]. The paper emphasizes that speedup increases with the number of proof branches.

The main wall-time results are summarized as follows:

| Setting | Native | Fallback | Speedup |
|---|---:|---:|---:|
| `mathd_numbertheory_345` | 132.8 s | 2,641.4 s | 19.9× |
| `mathd_numbertheory_3` | 116.2 s | 1,572.4 s | 13.5× |
| `mathd_algebra_478` | 119.9 s | 1,579.6 s | 13.2× |
| Holes 1 (5 probs), `B=7` | 196 s | 1,537 s | 7.9× (6.3–9.6×) |
| Holes 2 (19 probs), `B=14` | 136 s | 1,291 s | 9.4× (5.6–19.9×) |
| Holes 3 (11 probs), `B=21` | 131 s | 1,811 s | 13.8× (6.5–20.2×) |
| Holes 4 (7 probs), `B=28` | 125 s | 3,478 s | 24.6× (13.1–50.0×) |
| Holes 5 (3 probs), `B=35` | 141 s | 4,436 s | 29.8× (20.1–45.2×) |
| Average (all 45) | 140 s | 1,995 s | 14.0× |

The paper states that speedup grows *monotonically* with hole count $H$ and therefore with branch count $B$ [2605.25556]. Section 4.4 gives the scaling law:
$T_{\mathrm{native}}(C,H) \approx T_{\mathrm{elab}} + B \cdot T_{\mathrm{tactic}}$ and
$T_{\mathrm{fallback}}(C,H) \approx B \cdot (T_{\mathrm{load}} + T_{\mathrm{body}})$,
where $C = 7$ tactics and $B = H \cdot C$. On the reported hardware, $T_{\mathrm{elab}} \approx 120\ \mathrm{s}$ once per theorem, $T_{\mathrm{load}} + T_{\mathrm{body}} \approx 75\ \mathrm{s}$ per branch for simple theorems to approximately $795\ \mathrm{s}$ per branch for harder ones, and $T_{\mathrm{tactic}} \approx 0.045\ \mathrm{s}$ [2605.25556].

Projected runtimes extend the trend beyond the measured regime. For example, the paper reports projected values of approximately $8.7\times$ speedup at $B=14$, approximately $25.8\times$ at $B=42$, and approximately $34.3\times$ at $B=56$ [2605.25556]. At “classical DSP scale,” described as 100 draft candidates yielding, for example, $4$ holes $\times$ $7$ tactics $= 280$ branches, the paper estimates that fallback would take approximately 29 h on a single M-series laptop, whereas snapshots would reduce this to approximately 3.5 h. This suggests that the method is primarily a scaling intervention for portfolio search rather than a marginal constant-factor optimization.

## 6. Relation to import-level caching, limitations, and scope

The paper distinguishes proof-state snapshotting from import-level caching and presents the two as orthogonal techniques [2605.25556]. Import-level caching, designated “Level 1” and exemplified by Kimina Lean Server, eliminates only $T_{\mathrm{load}}$ and not $T_{\mathrm{body}}$. Its best per-branch speedup is given as $(T_{\mathrm{load}} + T_{\mathrm{body}})/T_{\mathrm{body}} \approx 1.1\times \ldots 5\times$ depending on $T_{\mathrm{body}}$. Snapshotting, designated “Level 2,” eliminates both $T_{\mathrm{load}}$ and $T_{\mathrm{body}}$ per branch, although it still pays $T_{\mathrm{load}}$ once per LSP server start.

The projected composition of the two levels is summarized in the paper as follows:

| Approach | Wall time | vs Level 0 |
|---|---:|---:|
| Level 0 (per-branch `lake`) | 2,641 s | 1× |
| Level 1 (import cached, `W=2`) | ~330 s | 8× |
| Level 2 (snapshot) | 133 s | 20× |
| Level 1+2 (persistent + snapshot) | ~17 s | >100× |

Several limitations are stated explicitly [2605.25556]. Accuracy is unchanged: snapshotting does not modify which goals are provable, only how fast tactics are tried. A patched binary is required, so deployment depends on trusting and installing a custom Lean 4 build via `elan`; without it, the system transparently falls back to `lake` builds. The current implementation restarts the LSP per theorem, so $T_{\mathrm{load}}$ is still paid once per theorem; a persistent-server variant would remove that one-time cost. Finally, all tests were conducted on an Apple M-series laptop, and the paper notes that other hardware and deeper proofs may shift the exact crossover points.

These caveats delimit the technique’s scope. Proof-state snapshotting is not presented as a new proving method, a new tactic portfolio, or a change to proof semantics. It is an execution-model modification for Lean 4 tactic search in which elaborated state is kept live and reused rather than repeatedly reconstructed. Within that scope, the reported results indicate that the dominant bottleneck in portfolio-based search on partially specified proofs can be moved from proof-state reconstruction to tactic execution itself [2605.25556].

Source: https://www.emergentmind.com/topics/proof-state-snapshotting