Topological Machine Learning
- Topological machine learning is a field that integrates algebraic topology techniques, like persistent homology and Mapper, to capture both global and local data structures.
- It employs tools such as persistence diagrams, persistence images, and the Euler Characteristic Transform to convert complex topological features into learnable representations.
- Recent advances incorporate differentiable topology layers and transformer architectures, enhancing applications in image analysis, phase detection, and materials discovery.
Topological machine learning is an emerging field that leverages techniques from algebraic topology to analyze complex data structures in ways that traditional machine learning methods may not capture. In the current literature, the term encompasses at least two closely related programs: the use of topological summaries such as persistent homology, persistence diagrams, Mapper graphs, and the Euler Characteristic Transform (ECT) inside conventional or deep learning pipelines, and the use of machine learning to identify, represent, or predict topological phenomena in physics and materials science, including topological phases, topological sectors, and topological defects (Coskunuzer et al., 2024, Rieck, 2024, Lian et al., 2019).
1. Conceptual scope and historical development
Applied topology in machine learning has evolved from an emphasis on the global shape of datasets to a broader treatment of local geometry. Early applications of persistent homology focused on dominant global structures, and short persistent homology bars were typically disregarded as sampling noise. More recent work argues that short bars can be highly discriminative for learning tasks because they encode local geometry, including texture, curvature, small-scale fluctuations, and even fractal dimension; the stated meta-hypothesis is that the short bars are as important as the long bars for many machine learning tasks (Adams et al., 2021).
Within this broader field, persistent homology and the Mapper algorithm are presented as two key techniques. Persistent homology captures multi-scale topological features such as clusters, loops, and voids, whereas the Mapper algorithm creates an interpretable graph summarizing high-dimensional data. The literature surveyed for practitioners treats these methods as part of a data-centric workflow spanning point clouds, images, graphs, and networks, and emphasizes their compatibility with conventional machine learning models after suitable representation or vectorization (Coskunuzer et al., 2024).
A distinct but related line of work supplies a topological formalism for supervised classification itself. Under very mild conditions, the classification problem in machine learning is described as always solvable, and a softmax classification network is interpreted as acting on an input topological space by a finite sequence of topological moves. In that formulation, architecture cannot be chosen independently from the shape of the underlying data; early dimension reduction can destroy separability for some manifold-structured datasets, whereas wider or higher-dimensional hidden layers provide more room to unknot, unlink, or glue labeled regions (Hajij et al., 2020).
2. Topological descriptors: persistent homology, persistence diagrams, and the Euler Characteristic Transform
Persistent homology begins from a filtration, namely a nested sequence of spaces. For each homological dimension , one records the birth and death of topological features across scale: The resulting persistence diagrams or barcodes separate long-lived features, often interpreted as robust global structure, from short-lived features, now frequently treated as informative local geometry rather than mere noise (Adams et al., 2021).
The Euler characteristic provides a different but related summary. For a -dimensional simplicial complex ,
where denotes the number of -simplices in . The ECT lifts this scalar invariant into a directional family of Euler characteristic curves. For an embedded simplicial complex , one defines a height function over directions , forms sublevel complexes 0, sets
1
and obtains
2
The overview centered on the ECT states that, with enough directions, the transform is injective, and uses this construction as a running example for more efficient models on point clouds, graphs, and meshes (Rieck, 2024).
For machine learning, both persistent homology and the ECT require representations compatible with standard models. Persistence diagrams are multisets of points in 3, and ECT outputs are functions indexed by direction and threshold. The standard practical move is therefore to convert these objects into fixed or learnable representations while preserving the stability and geometric fidelity that made the topological summary useful in the first place (Coskunuzer et al., 2024, Rieck, 2024).
3. Representations for learning: vectorizations, kernels, and architectures on topological summaries
A large part of topological machine learning concerns the representation of persistence diagrams. Surveyed vectorization methods include persistence landscapes, persistence images, persistence silhouettes, Betti curves, kernel methods, entropy summaries, representations using rings of algebraic functions, tropical coordinates, complex polynomials, and heat kernels. A recent classification pipeline makes this dependence explicit by treating the pair consisting of filtration and topological representation as the central design choice, and by using grid search to determine optimal representation methods and parameters for point clouds, images, and graphs (Adams et al., 2021, Conti et al., 2023).
Kernelization supplies one rigorous route from persistence diagrams to mainstream learning algorithms. The persistence scale-space kernel embeds diagrams into an 4-Hilbert space via a heat-equation construction, yields a positive definite kernel, and is proved stable with respect to the 5-Wasserstein distance. In the reported experiments on 3D shape classification, shape retrieval, and texture recognition, the proposed kernel outperformed the landscape kernel and was described as a theoretically sound connection between persistent homology and kernel SVMs or kernel PCA (Reininghaus et al., 2014).
A second route is functional compression. Persistence indicator functions map a persistence diagram to a step function that counts how many topological features are active at each parameter value. They can be calculated and compared in linear time, admit a parameter-free kernel-based similarity measure, and support statistical operations such as means and confidence bands that are difficult to define directly on diagrams (Rieck et al., 2019).
More recent work removes the requirement for hand-crafted vectorization altogether. Persformer is presented as the first Transformer neural network architecture that accepts persistence diagrams as input. It processes persistence diagrams as unordered sets, satisfies a universal approximation theorem, and outperforms previous topological neural network architectures on ORBIT5k, ORBIT100k, and MUTAG. Its saliency-based interpretability analysis also shows that some points close to the diagonal can have high saliency, reinforcing the broader claim that short bars may be predictive rather than disposable (Reinauer et al., 2021).
4. Differentiable topology and topology-aware deep learning
Topological machine learning is not limited to fixed preprocessing. A major development is the construction of differentiable topological modules that can be inserted directly into gradient-based pipelines. A topology layer for machine learning computes persistent homology based on level set filtrations and edge-based filtrations, and is used for three applications: regularizing data reconstruction or model weights, constructing a loss on the output of a deep generative network to incorporate topological priors, and performing topological adversarial attacks on deep networks trained with persistence features (Brüel-Gabrielsson et al., 2019).
The ECT has been pushed in the same direction. To make Euler-characteristic-based summaries compatible with end-to-end optimization, step functions in the Euler characteristic curve are replaced with sigmoids, producing a differentiable approximation. In that construction, gradients can flow through directions, thresholds, and even input coordinates, so the ECT can function as an inductive bias inside neural architectures rather than only as an external descriptor (Rieck, 2024).
The topological framework for deep learning gives a complementary interpretation of what such models are doing. Linear maps act as rotations, reflections, scalings, or quotienting; bias additions act as translations; nonlinearities such as ReLU can induce bending or quotienting; and softmax maps activations to the interior of a simplex, with classification determined by the corresponding Voronoi cell. On this view, a classifier learns a finite sequence of topological moves that continuously deforms the input space so that labeled subsets become separable, which also explains why width and early-layer dimensionality matter for manifold-structured data (Hajij et al., 2020).
5. Data modalities and empirical applications outside quantum topology
The principal empirical appeal of topological machine learning is its breadth across data types. The practitioner-oriented literature presents point clouds, images, graphs, and networks as canonical settings for persistent homology and Mapper, with case studies in shape recognition, cancer image analysis, genotyping, drug discovery, and anomaly detection in transaction networks (Coskunuzer et al., 2024). The classification pipeline based on filtration selection and persistence-diagram representation further reports benchmark results across modalities, including about 6 accuracy on a dynamical-systems point-cloud task, about 7 on MNIST using multiple image filtrations, about 8 on FMNIST, and about 9 on COLLAB graphs (Conti et al., 2023).
For multivariate time series, one proposed framework converts sliding windows into point clouds, computes persistent homology, compares persistence diagrams by Wasserstein distance, and performs supervised classification with 0-nearest neighbors. Because ordinary TDA is translation- and rotation-invariant and insensitive to coordinate semantics, the method introduces symmetry-breaking and anchor points for heterogeneous sensor variables. On room occupancy detection, the reported accuracies are 1 and 2 on two test sets, and on an activity-recognition dataset the reported test accuracy is 3 (Wu et al., 2019).
For mixed numeric and categorical data, TopMix addresses the lack of a direct point-cloud model by one-hot encoding categorical variables, standardizing all features, applying symmetry breaking, and constructing a point cloud for each data object using projection maps. Persistence diagrams are then compared with the 4-Wasserstein distance and classified with 5-nearest neighbors. On the Cleveland Heart Disease Dataset, the reported test accuracy is 6, and the method is described as the first general topological machine learning framework for mixed data (Wu et al., 2020).
The ECT literature adds a complementary perspective on efficiency. In point clouds, graphs, and meshes, discretized ECT matrices can be flattened into feature vectors or treated as images, and ECT-based features are presented as compact, fixed-size summaries that can reduce memory and computational requirements relative to raw high-dimensional input. For geometric graphs in particular, the overview states that ECT-based features can offer competitive performance compared with graph neural networks while requiring significantly less memory and computational overhead because the transform computes feature vectors via simple counts rather than message passing (Rieck, 2024).
6. Topological phases, defects, and materials discovery
A substantial part of the field concerns machine learning on systems whose targets are themselves topological. One early result shows that restricted Boltzmann machines can exactly and efficiently represent the one-dimensional symmetry-protected topological cluster state and the two-dimensional and three-dimensional toric code states, with the number of parameters scaling linearly with system size. The same framework represents excited states with abelian anyons and their mutual statistics, and reinforcement learning with an RBM variational ansatz captures a topological phase transition in a non-integrable Hamiltonian (Deng et al., 2016).
In experimental quantum simulation, a 7D convolutional neural network was trained on synthetic density matrices and then used to classify topological phases of a three-dimensional chiral topological insulator simulated in a nitrogen-vacancy center in diamond. Training and validation accuracy both approach about 8, misclassifications occur mostly near phase transitions, and the network still correctly identifies the topological phase with greater than 9 probability even when more than 0 of the density matrices are randomly removed (Lian et al., 2019).
Other architectures specialize to the geometry of topological quantum data. Quaternion-based unsupervised and supervised methods classify Chern insulators by transforming eigenstates or spin textures into quaternion-valued representations; the quaternion CNN is reported to achieve near-perfect accuracy up to 1 and to generalize to states with distributions different from those seen during training (Lin et al., 2022). In real-space condensed-matter models, eigenvector ensembling with decision trees and random forests recovers phase diagrams of Su-Schrieffer-Heeger systems and yields Shannon information entropy signatures that indicate how topological information is distributed in the bulk, which in turn motivates topological lattice compression (Holanda et al., 2019). In SU(3) Yang-Mills theory, convolutional neural networks estimate topological charge 2 from topological charge density at small flow time, but a dimension-reduction study finds no statistically significant dependence on input dimension, leading to the conclusion that the network relies on global integrated quantities rather than characteristic local structures (Kitazawa et al., 2019).
Topological materials discovery has become a major application area because large ab-initio databases now contain tens of thousands of labeled compounds. Gradient boosted trees trained on symmetry- and chemistry-based descriptors predict the topology of crystalline materials with an accuracy of 3 on 4 unique representatives, using features such as space group, number of electrons per unit cell, and coarse-grained orbital content (Claussen et al., 2019). A later study merges Materiae and the Topological Materials Database into a dataset of 5 materials, performs five-class and binary classification, and reports that XGBoost achieves 6 multiclass accuracy and 7 binary accuracy; maximum packing efficiency and the fraction of 8 valence electrons are highlighted as critical features (He et al., 20 Mar 2025). For two-dimensional topological insulators, machine-learning-accelerated screening based on atomic and prototype features is reported to determine electronic topology with an accuracy of over 9, to discover 0 non-trivial materials including 1 novel insulating candidates corroborated by density functional theory, and to be 2 more efficient than trial-and-error search (Schleder et al., 2021).
Topological defect formation extends the scope of the field from classification to dynamical prediction. Using a recurrent neural network, one study shows that the final spatial configuration of topological defects created during a second-order phase transition can be predicted from the time evolution of the order parameter over a short interval near the critical point, well before equilibration. The same work reports that the predictability of the machine-learning model follows the power-law scaling dictated by the Kibble-Zurek mechanism (Suzuki et al., 28 Aug 2025).
7. Misconceptions, interpretability, and future directions
One recurring misconception is that only the most persistent topological features matter. The survey literature explicitly argues against this by emphasizing the discriminative role of short bars in local geometry (Adams et al., 2021). Persformer’s saliency analysis reaches a similar conclusion from a learned model: large-persistence points are often salient, but some points near the diagonal can also have high saliency, and on curvature regression datasets the shortest bars can be especially salient (Reinauer et al., 2021). These results do not eliminate the classical distinction between robust global features and near-diagonal features, but they do narrow the range of tasks for which “short bars = noise” is an adequate summary.
A second misconception is that topological machine learning always extracts genuinely local topological mechanisms from high-dimensional input. The Yang-Mills study provides a counterexample: although the network receives high-dimensional topological charge density fields, its accuracy is statistically independent of input dimension after dimensional reduction, suggesting that its successful predictions come from global integrated quantities rather than higher-dimensional local structure (Kitazawa et al., 2019). This indicates that interpretability remains an open methodological issue even when the prediction target is topological.
Current forward-looking accounts describe three directions for the field: the learning of functions on topological spaces, the building of hybrid models that imbue neural networks with knowledge about the topological information in data, and the analysis of qualitative properties of neural networks (Rieck, 2024). Together with practitioner-oriented tutorials that emphasize hands-on integration of persistent homology and Mapper into standard workflows (Coskunuzer et al., 2024), these directions suggest a convergence between topology-aware representations, differentiable topological layers, and domain-specific inductive biases. A plausible implication is that future progress will depend less on treating topology as a single feature extractor and more on matching the choice of filtration, representation, differentiable module, and architectural constraints to the geometry of the data and to the topology of the prediction target.