---
title: MLFMF Dataset Overview
url: https://www.emergentmind.com/topics/mlfmf-dataset
type: topic
---

# MLFMF Dataset Overview

MLFMF is a collection of datasets designed for benchmarking machine learning methods in the formalization of mathematics, targeting recommendation systems for proof assistants. These datasets support rigorous evaluation of premise selection, reference prediction, and related learning tasks based on structured libraries of formalized mathematics from major proof assistant ecosystems. MLFMF leverages both syntactic and graph-theoretic representations—incorporating heterogeneous dependency graphs and full abstract syntax trees—to enable graph neural models, structure-aware embeddings, and conventional NLP approaches. It is currently the largest corpus of machine-learnable formalized mathematical knowledge, encompassing over $250{,}000$ entries extracted from Agda and Lean 4 libraries [2310.16005].

## 1. Scope and Library Composition

MLFMF integrates four principal mathematical libraries reflecting diverse proof assistant paradigms, formalization styles, and mathematical subdomains:

- **Lean 4 Mathlib**: The primary, rapidly evolving Lean 4 mathematics library, noted for extensive tactic use, a broad subject spectrum, and an active premise-selection research community.
- **Agda Standard Library**: The core Agda library, characterized by tactic-free, “full-term” proofs and coverage of algebra, combinatorics, and order theory.
- **Agda-unimath**: A large Agda corpus formalizing univalent mathematics and Homotopy Type Theory, with emphasis on higher-dimensional algebraic structures.
- **TypeTopology (Agda)**: A mid-sized Agda sublibrary focusing on formal topology, featuring technical constructions and intricate dependency chains.

Each library is parsed and represented in two modalities: (1) as a directed heterogeneous network encoding the modular structure and cross-references among entries, and (2) as a per-entry s-expression dump of its elaborated abstract syntax tree (AST).

## 2. Data Extraction and Representation

### Heterogeneous Network

For each library, MLFMF constructs a directed graph $G=(V,E)$, where nodes $V$ include:

- Library node (one per corpus)
- Module/namespace nodes (file-based for Agda, declared for Lean)
- Entry nodes (such as functions, data types, records, axioms, etc.), each labeled with entry kind (e.g., `:data`, `:function`, `:axiom`)

Edges $E$ encode:

- **CONTAINS**: library $\rightarrow$ module, module $\rightarrow$ submodule
- **DEFINES**: module $\rightarrow$ entry
- **REFERENCE_FROM_DECLARATION**: entry $u$ references entry $v$ in its type signature
- **REFERENCE_FROM_BODY**: entry $u$ references entry $v$ in its implementation or proof term

This network provides a granular map of the formal library's modular and referential architecture, enabling graph-based relational learning.

### S-expression AST Dump

For each entry, the fully elaborated AST is exported as a Lisp-style s-expression, structured as $(name~decl~body)$ where:

- `name`: fully-qualified identifier
- `decl`: s-expression encoding the type, with implicitness and binding details
- `body`: s-expression for the proof term or definition

In Lean, shared subterms lead to a DAG structure; in Agda, trees are true syntactic trees. Each entry line is independent and fully self-contained, permitting direct batch parsing and analysis.

## 3. Statistical Overview

Dataset statistics across all four libraries are as follows:

| Library         | Entries       | Nodes         | Edges         |
|-----------------|--------------|--------------|--------------|
| Agda Stdlib     | $\approx38{,}000$ | $\approx40{,}500$ | $\approx95{,}000$  |
| Agda-unimath    | $\approx79{,}000$ | $\approx82{,}500$ | $\approx190{,}000$ |
| TypeTopology    | $\approx12{,}500$ | $\approx13{,}000$ | $\approx30{,}000$  |
| Lean Mathlib4   | $\approx120{,}000$ | $\approx123{,}500$ | $\approx300{,}000$ |

Across all corpora, the total number of entries is $\sum |\mathcal{T}| \approx 250{,}000$. Mean per-entry (tree or DAG) node count is approximately 420, with standard deviation 200; average AST depth is $\approx 15$ (std $\approx 7$); and the vocabulary size for s-expression tokens is approximately 10,000.

## 4. File Formats and Organization

Each library’s data is distributed in a canonical folder structure:

```
mlfmf/
  agda/
    stdlib/
      network.graphml    # Heterogeneous network in GraphML
      entries.sexpr      # One s-expression per entry
    unimath/
      ...
    typetopo/
      ...
  lean/
    mathlib4/
      network.graphml
      entries.sexpr
```

**network.graphml** files encode the node and edge types for $G(V,E)$.  
**entries.sexpr** are plain UTF-8 text files, one s-expression per line, directly parseable with standard Lisp parsers. Typical loading code involves common Python libraries:

```python
import networkx as nx
G = nx.read_graphml('network.graphml')

from sexpdata import loads
with open('entries.sexpr') as f:
    for line in f:
        name, decl, body = loads(line)
```

This structure supports direct integration with graph analytics, NLP pipelines, and structure-aware ML models.

## 5. Benchmarking Protocols and Baseline Results

### Tasks

Two primary tasks are benchmarked:

- **Link prediction**: Predicting presence of missing REFERENCE edges in $G$—paradigmatic for premise selection.
- **Recommendation**: Top-$k$ prediction of missing dependencies in the context of an incomplete proof construction.

**Train/test protocol**: $p_{\mathrm{test}} = 0.2$ of function nodes are “masked” as incomplete; within each, $p_{\mathrm{body}} = 0.1$ of reference edges are removed (as positives), with negatives sampled from non-edge pairs.

### Baseline Methods

1. **Dummy**: Ranks entries by static in-degree.
2. **BoW Jaccard**: Token set overlap for entry s-exprs, $J(A,B) = \frac{|A \cap B|}{|A \cup B|}$.
3. **TF-IDF**: Similarity by cosine and Manhattan distances in TF-IDF-weighted token space.
4. **fastText**: Pretrained embedding averages (Common Crawl), weighted by TF-IDF.
5. **Analogy-based**: fastText vector offsets.
6. **node2vec + tree bagging**: node2vec embeddings on $G$ followed by a “tree-bagging” (ensemble of 100 trees) classifier.

### Summary of Results

| Method         | Agda stdlib acc/minRank | Agda unimath acc/minRank | TypeTopology acc/minRank | Lean Mathlib4 acc/minRank | 
|----------------|-----------------------|-------------------------|-------------------------|--------------------------|
| Dummy          | 0.51 / 218            | 0.53 / 2134             | 0.50 / 4556             | 0.51 / 26065             |
| BoW            | 0.50 / 1608           | 0.50 / 1571             | 0.50 / 4496             | 0.50 / 15458             |
| TF-IDF         | 0.51 / 144            | 0.52 / 112              | 0.51 / 552              | 0.51 / 443               |
| fastText       | 0.51 / 132            | 0.52 / 394              | 0.50 / 1292             | NA                       |
| Analogies      | 0.52 / 37             | 0.51 / 158              | NA                      | NA                       |
| node2vec       | 0.96 / 4.4            | 0.96 / 3.2              | 0.98 / 5.8              | 0.95 / 195               |

**Interpretation**: node2vec-based approaches substantially outperform token-based or in-degree baselines (mean minimal rank $<6$ for Agda), whereas text-centric models yield marginal improvements over the dummy. Lean Mathlib4 remains challenging for minimal-rank recommendation due to tactic-heavy idioms and sparser explicit dependency tagging.

## 6. Use Cases and Extensibility

MLFMF supports a range of tasks in mathematical machine learning for formalization, including:

- **Premise selection** via link-prediction and recommendation.
- **Graph neural network training** over $G(V,E)$ with per-node and per-edge features from s-exprs.
- **Joint text-structure models** that integrate AST sequence data and dependency graph architecture.
- **Entry classification** (e.g., distinguishing functions from axioms).
- **Transfer learning** across proof assistant frameworks and formalization styles.

The extraction pipeline generalizes: for a new proof assistant, a “lib2sexp” extension or plugin can generate s-expr dumps; a “sexp2graph” parser constructs the network. The modular pipeline is agnostic to the underlying logic, allowing extension to Coq, Isabelle/HOL, HOL Light, Metamath, or similar systems with modest implementation effort.

## 7. Significance in Formalized Mathematics and Machine Learning

MLFMF provides the largest machine-learnable corpus of formalized mathematics to date, spanning over a quarter-million entries with semantic and syntactic structure. Its dual representations—heterogeneous networks and AST s-expressions—render it uniquely suitable for evaluating premise selection, reference prediction, and proof automation techniques that require deep structural and linguistic understanding. The dataset has become a standard for empirical validation in machine learning for mathematical formalization and is poised for further expansion into new proof assistant ecosystems [2310.16005].

Source: https://www.emergentmind.com/topics/mlfmf-dataset