---
title: Matrix Product State Classifiers
url: https://www.emergentmind.com/topics/matrix-product-state-mps-classifiers
type: topic
---

# Matrix Product State Classifiers

Matrix Product State (MPS) classifiers are classification methods that use the matrix product state ansatz—equivalently the tensor-train format—as the central representation for inputs, model parameters, or quantum feature states. In the literature, the term covers several distinct constructions: an MPS can serve as an unsupervised tensor-compression front end whose low-order core is passed to a conventional classifier, it can be the trainable classifier itself in classical or quantum-inspired form, or it can act as an encoding or simulation backend for another classifier. Across these settings, the common premise is that high-order data or high-dimensional feature maps can often be represented with moderate bond dimension when the relevant correlations are sufficiently structured [1503.00516] [1609.04541] [1905.01426] [2412.15826].

## 1. Scope and meanings of “MPS classifier”

The literature covered here shows that the label “MPS classifier” does not denote a single model class. In one line of work, MPS is a multilinear feature extractor: training samples are concatenated into a higher-order tensor, decomposed by successive singular value decompositions (SVDs), and replaced by compact core features that are then classified by methods such as **KNN-1** or **LDA**. In another line, the MPS itself is the supervised model, producing class scores or probabilities from an embedded input. In a third line, MPS is neither the classifier nor the learned decision rule, but the computational mechanism that makes a quantum kernel or a compressed quantum input representation tractable [1503.00516] [1905.01426] [2411.09336].

| Usage | Role of MPS | Representative papers |
|---|---|---|
| Feature-extraction front end | Learns common factors and low-order core features, then uses standard classifiers | [1503.00516], [1609.04541] |
| Direct trainable classifier | Produces class scores or probabilities from embedded inputs | [1905.01426], [2406.17441], [2412.15826] |
| Encoding or simulation backend | Compresses inputs or simulates quantum feature states for another classifier | [2406.06935], [2411.09336] |

A persistent misconception is that all “MPS classifiers” are supervised tensor-network decision functions. That is not the case. The early tensor-classification papers built the classification stage from conventional models after MPS compression, and they explicitly state that labels are not used during the MPS decomposition itself [1503.00516] [1609.04541]. Conversely, some later works use the phrase for explicitly trainable MPS models, including quantum-circuit realizations and probabilistic generative classifiers [1905.01426] [2412.15826].

## 2. MPS as tensor feature extraction for classification

A foundational use of MPS in classification is as a compression mechanism for tensor-valued data of order \(N\ge 3\). Given training samples
\[
\bm{X}^{(k)}\in \mathbb{R}^{I_1\times I_2\times \cdots \times I_N},\quad k=1,\dots,K,
\]
the training set is stacked into a tensor, the sample mode is placed near the center of the chain, and a mixed-canonical MPS decomposition is computed by a left-to-right and right-to-left sequence of truncated SVDs. The samplewise representation is
\[
x_{i_1\cdots k\cdots i_N} \approx B^{(1)}_{i_1}\cdots B^{(n-1)}_{i_{n-1}}\, G_k^{(n)}\, C^{(n+1)}_{i_n}\cdots C^{(N+1)}_{i_N},
\]
where the learned \(B\) and \(C\) tensors are common factors and \(G_k^{(n)}\in\mathbb{R}^{\Delta_{n-1}\times \Delta_n}\) is the core matrix for sample \(k\). This reduces each sample to
\[
N_f=\Delta_{n-1}\Delta_n
\]
features, with a core tensor of maximum order three regardless of the order of the original data [1503.00516] [1609.04541].

This construction was introduced as an alternative to Tucker/HOOI feature extraction. The stated motivations are structural and computational. Structurally, Tucker retains a high-order core, whereas MPS always reduces the extracted feature object to a matrix or third-order tensor. Computationally, MPS is obtained by successive truncated SVDs, not by alternating least squares. Under the simplifying assumption \(I_j\equiv I\), \(\Delta_j\equiv \Delta\), the per-iteration HOOI complexity is reported as
\[
O\!\left(KI^N + NKI^{2(N-1)} + NK^{3(N-1)}\right),
\]
with HOSVD initialization cost \(O(NKI^{N+1})\), while the dominant MPS cost is
\[
O(KI^{N+1}),
\]
dominated by the first SVD [1609.04541].

The practical pipeline is unsupervised up to the final classifier. Training samples are concatenated, the sample mode is permuted into position \(n\), MPS factors and training cores are extracted with an SVD threshold
\[
\frac{\sum_{i=1}^{\Delta_j} s_i}{\sum_{i=1}^{r_j} s_i}\ge \epsilon,
\]
test tensors are projected with the learned common factors, and the resulting cores are vectorized before classification by **KNN-1** or **LDA** [1503.00516] [1609.04541].

Empirically, these papers report that MPS often attains higher **classification success rate (CSR)** with fewer features than HOOI. On **COIL-100**, one best reported result at \(r=50\%\) is **\(99.19\pm0.19\)** for MPS with \(N_f=120\) and \(\epsilon=0.80\), versus **\(98.87\pm0.19\)** for HOOI with \(N_f=198\). On **Extended Yale Face Database B**, MPS reaches **\(97.32\pm0.89\%\)** with **LDA**, compared with roughly **\(95.89\%-96.07\%\)** for HOOI. On **BCI Jiaotong**, **Subject 2** reaches **\(91.02\pm0.70\%\)** for MPS versus about **\(79\%-83\%\)** for HOOI [1609.04541]. These studies also note an important caveat: on some datasets, even the MPS core remains too large for direct classification, so an additional empirical truncation of core dimensions is applied before the final classifier [1503.00516] [1609.04541].

## 3. Direct supervised MPS classifiers and quantum-circuit realizations

A more literal use of the term is the direct trainable MPS classifier. One quantum-circuit realization defines a binary dataset \(S=\{(x^d,y^d)\}_{d=1}^D\), normalizes features to \([-\pi,\pi]\), and embeds feature \(x_n^d\) as
\[
\phi_n^d=\cos(x_n^d)\ket{0}+\sin(x_n^d)\ket{1}.
\]
The full input is the product state \(\phi^d=\otimes_{n=1}^N \phi_n^d\). A sequential circuit of unitary blocks built from single-qubit rotations around the \(y\)-axis and **CNOT** gates processes the chain, discarding one qubit from each block and propagating the remaining qubit forward until a final output qubit is measured. Training minimizes
\[
J_{\theta}=\dfrac{1}{D}\sum_{d=1}^{D}(M_{\theta}(x^d)-y^d)^2
\]
with **conjugate gradient** optimization [1905.01426].

That model was evaluated on pairwise binary tasks derived from **Iris** and a meteorological **Agri/ET\(_o\)** dataset on **ibmqx4**. Reported test accuracies include **90** for **Iris\(_3\)** and **80.65** for **Agri\(_1\)**. The paper frames bond dimension as the parameter controlling expressivity and warns that “larger dimension of bond results in higher accuracy,” but that an “extremely large bond dimension” can also lead to overfitting [1905.01426].

A different direct formulation uses a supervised classical MPS whose output amplitudes are squared to obtain class probabilities. Two architectures are described: an ensemble with one MPS per class, and a single MPS with a central label tensor. In the ensemble formulation,
\[
p(c=i\mid x)=\frac{1}{Z}y_i^2,\qquad Z=\sum_c y_c^2.
\]
Training uses **cross-entropy loss**, automatic differentiation, and identity-like initialization
\[
A_s = \frac{I_d}{\sqrt{d}} + \hat{A}^s,
\qquad
\hat{A}^s \sim \mathcal{N}(0,\sigma^2),
\]
rather than DMRG-style sweeping. This same supervised MPS is then repurposed as a generator through class-conditional sampling from reduced density matrices, and a GAN-style training stage is used to improve sample realism by reducing outliers while preserving classification accuracy [2406.17441].

The resulting picture is that a direct MPS classifier need not be purely discriminative. In this literature, the same MPS can function as a supervised model, a class-conditional density model, and a sequential generator, provided the local embedding supports tractable marginalization [2406.17441].

## 4. Probabilistic MPS classifiers for time-series data

Time series provide a natural one-dimensional domain for MPS classifiers. In **MPSTime**, a univariate series \(\mathbf{x}=(x_1,\dots,x_T)\) is mapped to a product feature state
\[
\Phi(\mathbf{x}) = \phi_1(x_1)\otimes \phi_2(x_2)\otimes \cdots \otimes \phi_T(x_T),
\]
with \(\phi_t(x_t)=[b_1(x_t),\dots,b_d(x_t)]\), using either orthonormal **Legendre** or **Complex Fourier** basis functions. For classification, the class label is attached as an index to one MPS tensor, yielding class-conditional amplitudes
\[
f^l(\mathbf{x}) = W^l \cdot \Phi(\mathbf{x}),
\qquad
p(\mathbf{x}\mid l)=|f^l(\mathbf{x})|^2,
\]
and the prediction rule
\[
\hat l = \operatorname*{arg\,max}_l |f^l(\mathbf{x})|^2.
\]
This makes the classifier generative for classification: it models the class-conditional joint distribution over the entire time series rather than only a discriminative boundary [2412.15826].

Training uses negative log-likelihood, a DMRG-inspired sweeping algorithm, and a local **tangent-space gradient optimization (TSGO)** update on merged two-site bond tensors,
\[
\mathcal{B}' = \mathcal{B} - \eta \frac{\partial \mathcal{L}/\partial \mathcal{B}}{\left|\partial \mathcal{L}/\partial \mathcal{B}\right|},
\]
followed by normalization, SVD splitting, and truncation to \(\chi_{\rm max}\). For classification, preprocessing combines a scaled outlier-robust sigmoid transform with min-max normalization to \([-1,1]\) [2412.15826].

On **ECG200**, **ItalyPowerDemand**, and an **Astronomy** benchmark based on a balanced Kepler subset, MPSTime is reported as competitive with **InceptionTime** and **HIVE-COTE 2.0**, while outperforming **1-NN-DTW** on ECG and Power Demand. The paper prints an explicit **Astronomy** mean accuracy of
\[
0.91 \pm 0.01,
\]
compared with **\(0.93 \pm 0.01\)** for InceptionTime and **\(0.96 \pm 0.01\)** for HC2 [2412.15826]. A notable modeling claim is that classification requires smaller physical dimension \(d\) than imputation: for ECG and Power Demand, optimal values were \(d=3\text{--}6\) and \(\chi_{\rm max}=15\text{--}25\), whereas the broader application range reported for MPSTime is \(\chi_{\rm max}=20\text{--}160\) [2412.15826].

## 5. Ordering, locality, and architectural modifications

A recurrent theme in MPS classifiers is that ordering is not a neutral implementation detail. In amplitude-encoded quantum inputs, the truncation loss of an MPS approximation depends not only on the bond dimension \(\chi\), but also on how classical features are mapped onto qubits. One work therefore searches over qubit permutations with **uniform-cost search**, using cumulative truncation error as path cost. The central claim is that optimized feature-to-qubit mapping reduces Frobenius reconstruction loss and improves downstream classifiers, including an **MPS classifier with bond dimension 10**, **Adam**, learning rate **\(10^{-2}\)**, batch size **128**, and **300** training epochs. On **MNIST**, **Fashion-MNIST**, and preprocessed **CIFAR-10**, classifiers trained on permuted MPS encodings outperform those trained on standard encodings, with the largest gains at low input bond dimension; the paper also reports that the MPS classifier trained with permuted MPS images converges more quickly [2406.06935].

This sensitivity to ordering is closely related to a structural limitation of vanilla MPS: the exponential decay of correlations along the chosen one-dimensional chain. For flattened images and other non-sequential data, nearby semantic variables may become distant in chain order. **Shortcut Matrix Product States (SMPS)** address this by adding long-range shortcut bonds between selected tensors. The paper’s architectural claim is that shortcuts can “decrease significantly the correlation length of the MPS” while preserving much of its tractability. Although the main experiments concern function fitting, partition function calculation, and unsupervised generative modeling rather than discriminative classification, the paper explicitly situates the work against the weakness of MPS on long-range dependencies and discusses supervised kernel linear classification as a motivating use case [1812.05248].

A related input-side result shows that transform choice can materially change MPS compressibility. By applying a **discrete wavelet transform (DWT)** before MPS compression, one study prepares a **\(128\times128\)** ChestMNIST image on **14 qubits** with fidelity exceeding **99.1%** on a circuit with total depth **425** single-qubit rotations and CNOT gates. This is not a classifier result, but it suggests that multiscale preprocessing can expose low-entanglement structure before an MPS-based classifier or quantum encoding stage is applied [2502.16464].

## 6. Boundaries, misconceptions, and recurrent limitations

Not every classifier that involves MPS is an MPS classifier in the direct-model sense. A large-scale quantum kernel study uses MPS only as a simulator for quantum feature states. The learned model is a classical **SVM** with Gram matrix
\[
\boldsymbol{K}_{ij} = \left|\langle \psi(\boldsymbol{x}_i),\psi(\boldsymbol{x}_j)\rangle\right|^2,
\]
while MPS serves as the backend that makes state preparation and overlaps tractable. The paper reaches **165 features** and **6400 training data points**, but it is explicit that this is not a direct tensor-network classifier; it is a quantum kernel classifier whose states are simulated with MPS [2411.09336].

A further terminological boundary concerns uses of “classification” outside machine learning. One paper classifies one-dimensional gapped phases protected by **matrix product operator (MPO) symmetries** through local \(L\)-symbols satisfying coupled pentagon equations. This is MPS-based phase classification, not supervised prediction from labeled data [2203.12563]. Likewise, efficient algorithms for learning the closest MPS representation of a quantum state concern tomography and model recovery rather than classification, even though they are relevant to the broader question of how compact MPS descriptions can be reconstructed [2510.07798].

Within machine learning proper, several limitations recur across the literature. Feature-extraction pipelines based on MPS are unsupervised with respect to labels and sometimes require additional empirical truncation of the core before classification [1503.00516] [1609.04541]. Direct trainable models inherit the one-dimensional inductive bias of the chain, so mode ordering, core position, and locality assumptions can strongly affect performance [2406.06935] [1812.05248]. MPS is also not uniformly dominant: **TTPCA** or **CSP** can outperform it on some tasks, and stronger entangling structure does not necessarily improve generalization in MPS-backed quantum models [1609.04541] [2411.09336]. Even claims of “global optimality” in SVD-based MPS compression are explicitly narrower than global optimization of the full nonconvex tensor problem: the guarantee is that each matrix truncation step is SVD-optimal for the fixed ordering, not that every tensor approximation objective is solved globally [1609.04541].

Taken together, these works define MPS classifiers less as a single algorithm than as a research area organized around one principle: a classifier can be built around a one-dimensional low-rank tensor-network representation of data, weights, or quantum feature states. The main technical questions then become how to place variables along the chain, how much bond dimension is needed, whether labels enter through a separate classifier or through the MPS itself, and how much of the task-relevant structure survives the compression.

Source: https://www.emergentmind.com/topics/matrix-product-state-mps-classifiers