Quantum Extreme Learning Machines
- Quantum Extreme Learning Machines (QELMs) are fixed quantum-feature-map models where the quantum substrate encodes inputs and only the output layer is trained via linear regression.
- They are implemented on diverse platforms, including superconducting circuits and photonic systems, and have been applied in quantum state estimation, entanglement witnessing, and time-series prediction.
- Challenges such as classical simulability, concentration phenomena, and noise sensitivity drive research toward co-designing encoding strategies, reservoir dynamics, and measurement protocols.
Quantum Extreme Learning Machines (QELMs) are quantum variants of Extreme Learning Machines in which a fixed quantum substrate realizes the hidden-layer feature map and only the output layer is trained, typically through linear regression or a related single-layer classical model (Mujal et al., 2021, Lorenzis et al., 8 Sep 2025). In the strict definition used in the review literature, QELMs are memoryless: the quantum substrate is reinitialized between inputs, so each input instance is processed independently and the output depends only on the current input (Mujal et al., 2021). Recent work has extended this paradigm from static classical classification and regression to quantum state estimation, entanglement witnessing, open-system dynamics characterization, molecular modeling, exoplanetary retrieval, software testing, and time-series prediction, while also developing formal theories of effective measurements, Fourier expressivity, Pauli-transfer-matrix structure, concentration phenomena, and resource scaling (Innocenti et al., 2022, Xiong et al., 2023, Gross et al., 20 Feb 2026).
1. Canonical architecture and operational model
The standard QELM pipeline is hybrid quantum-classical. An input, which may be classical data or a quantum state, is encoded into a quantum system; the encoded state evolves under a fixed quantum channel or Hamiltonian; selected observables are measured; and the resulting measurement statistics are processed by a trained linear readout (Mujal et al., 2021, Innocenti et al., 2022). A common formalization writes the measured probability vector as
followed by a prediction
with obtained from linear regression or a pseudoinverse solution (Innocenti et al., 2022).
This architecture admits multiple realizations. In quantum-input tasks, the input can be a set of quantum states , the reservoir can be a unitary or quantum channel , and the feature matrix can be formed from expectation values
which are then mapped to targets by a linear readout (Vetrano et al., 2024). In classical-data tasks, representative encodings include dense angle encoding, Fourier or rotation encoding, quadrature displacements in continuous-variable photonics, and frequency-bin or spatial phase encoding in photonic platforms (Lorenzis et al., 8 Sep 2025, Monaco et al., 2024, Maier et al., 15 Oct 2025, Brusaschi et al., 20 Mar 2026, Joly et al., 16 May 2025).
Several concrete reservoir models recur across the literature. One line uses spin or qubit networks evolving under XX, transverse-field Ising, or kicked Ising Hamiltonians (Lorenzis et al., 8 Sep 2025, Assil et al., 17 Mar 2026, Dao et al., 13 Mar 2026). Another uses Gaussian photonic substrates or Gaussian Boson Sampling (GBS) interferometers (Maier et al., 15 Oct 2025, Montesinos et al., 13 Jun 2026). Photonic random layers have also been realized through orbital angular momentum quantum walks, electro-optic modulation in frequency bins, and multimode fibers (Zia et al., 25 Feb 2025, Brusaschi et al., 20 Mar 2026, Joly et al., 16 May 2025). Across these implementations, the defining operational feature remains the same: the quantum part is fixed and untrained, while the classical readout is the only learned component (Mujal et al., 2021).
| Component | Typical role | Representative realizations |
|---|---|---|
| Input encoding | Map classical or quantum input into a quantum state | Dense angle encoding, Fourier encoding, quadrature displacements, frequency-bin encoding |
| Reservoir | Fixed feature map generated by quantum dynamics | XX Hamiltonian, transverse-field Ising, kicked Ising, Gaussian photonic circuit, multimode fiber |
| Readout | Train only the output layer | Moore-Penrose pseudoinverse, ridge regression, logistic regression, single-layer classifier |
The same architectural template also clarifies the relation between QELM and QRC. QRC exploits memory and is suited for temporal tasks, whereas QELM, in its strict definition, has no memory and is suited for non-temporal tasks such as classification and regression (Mujal et al., 2021). Later work has partially relaxed this separation by introducing temporal augmentation into QELM feature spaces, but the canonical definition remains memoryless.
2. Effective measurements, feature expressivity, and interpretability
A central theoretical development is the effective-measurement view. For a QELM specified by a channel and measurement , the observed statistics can be re-expressed as a measurement performed directly on the input state through effective POVM elements
The prediction then becomes
This leads to a sharp characterization: a QELM can reproduce 0 if and only if 1 lies in the real span of the effective POVM elements (Innocenti et al., 2022). Full quantum state tomography is therefore possible only when the effective POVM is informationally complete; otherwise, only expectations within that span are accessible (Innocenti et al., 2022).
This framework also formalizes a key limitation. QELMs can only retrieve information that is linearly accessible from the measurement statistics, so nonlinear functionals such as purity require additional copies of the input state or multiple-injection protocols (Innocenti et al., 2022). Distributed architectures extend this observation by showing that nonlinear targets such as polynomial functions, Rényi entropies, concurrence, and negativity can be reconstructed by using multiple copies and either sequential multiple injection or spatially multiplexed subsystems with entangling interactions (Gili et al., 12 Feb 2026). For linear tasks, spatial multiplexing yields a linear reduction in resource requirements per unit; for nonlinear tasks, the distributed design reduces per-reservoir requirements by replacing growth of a single reservoir with growth in the number of interacting subsystems (Gili et al., 12 Feb 2026).
A complementary line of theory analyzes QELM expressivity through Fourier decomposition. In this view, the prediction function can be written as a Fourier series whose frequencies are determined by the data encoding scheme, while the coefficients depend on both the reservoir and the measurement (Xiong et al., 2023). The number of accessible frequencies and the number of observables impose a formal upper bound on expressivity,
2
so increasing reservoir complexity alone does not remove encoding or measurement bottlenecks (Xiong et al., 2023). The practical consequence is that the encoding fixes the function class available to the model, while the reservoir and readout determine which parts of that function class are actually usable.
The Pauli-transfer-matrix (PTM) formalism makes the same conclusion explicit in operator language. In n-qubit QELMs with continuous-time reservoir dynamics, the encoding determines the complete set of nonlinear features available to the QELM, while the quantum channels linearly transform these features before they are probed by the chosen measurement operators (Gross et al., 20 Feb 2026). Optimizing a QELM can therefore be cast as a decoding problem in which one shapes the channel-induced transformations such that task-relevant features become available to the regressor (Gross et al., 20 Feb 2026). This interpretability perspective identifies a classical representation of the QELM and clarifies that the quantum reservoir does not create nonlinearity by itself; rather, it mixes, spreads, or filters features generated by encoding.
3. Dynamics, memory, and the boundary with reservoir computing
The strict memoryless definition of QELM has been challenged by recent architectures that import temporal structure into the feature space. In the characterization of non-Markovian open-system dynamics, a QELM can process quantum states 3 generated by successive system-environment collisions, with a disordered many-body quantum system as reservoir and local observables
4
as features (Assil et al., 17 Mar 2026). The study compares a baseline protocol using the instantaneous feature vector 5, an observable-augmented protocol that adds same-time measurements in other bases, and a memory-enhanced protocol that concatenates temporal features,
6
where 7 or a distant step such as the initial state (Assil et al., 17 Mar 2026).
The reported empirical hierarchy is
8
with the gap increasing as channel memory increases (Assil et al., 17 Mar 2026). Temporal extensions of the feature vector consistently and significantly enhance estimation accuracy relative to the baseline protocol; incorporating memory from earlier time steps yields the most substantial and robust improvements; and extensions based solely on additional observables offer only marginal gains (Assil et al., 17 Mar 2026). In the Markovian regime 9, all protocols perform well and are similar, whereas in the strongly non-Markovian regime the temporal extension becomes indispensable for high-fidelity estimation (Assil et al., 17 Mar 2026). This suggests that environmental memory effects can serve as a constructive resource for learning.
A related development is the time-delayed QELM (TD-QELM) for time-series prediction on NISQ hardware. TD-QELM encodes multiple past inputs simultaneously into different qubits, achieving circuit depth independent of sequence length and reducing overall quantum computational cost from 0 in standard QRC to 1 (Kawanabe et al., 25 Feb 2026). On the NARMA benchmark, the best reported TD-QELM error is 2, versus 3 for QRC, and TD-QELM remains low and flat as the input length grows to 4 (Kawanabe et al., 25 Feb 2026). These results place TD-QELM in an intermediate zone between classical time-delayed ELMs and quantum reservoir computing.
Dynamics also matter in static quantum inference. In QELM-based quantum state estimation, accurate reconstruction remains possible even beyond the scrambling time, and at long interaction times the reconstruction efficiency matches the optimal one offered by random global unitary dynamics for all cases studied (Vetrano et al., 2024). The reported long-time values converge empirically to 5 for MSE, 6 for local Holevo information, and near-maximal OTOC saturation, while numerical conditioning remains manageable (Vetrano et al., 2024). This reinterprets scrambling: local information may be delocalized, yet still sufficient for QELM reconstruction.
4. Physical platforms and implementation modalities
The QELM literature spans a wide range of quantum hardware. The review literature lists spin or qubit networks, Fermi-Hubbard and Bose-Hubbard systems, networks of harmonic oscillators, continuous-variable systems, NMR, trapped ions, superconducting circuits, and Gaussian oscillator networks (Mujal et al., 2021). Later work has turned many of these into concrete task-specific implementations.
| Platform class | Representative implementations | Reported point |
|---|---|---|
| Superconducting and digital qubit processors | IBM Brisbane for molecular PES/FFs, IBM Fez for exoplanet retrieval, ibm_kawasaki for TD-QELM, ibm_quebec for large-scale kicked-Ising QELM | Real-hardware demonstrations from 4–7 qubits up to 124 qubits and more than 5,000 two-qubit gates (Monaco et al., 2024, Vetrano et al., 3 Sep 2025, Kawanabe et al., 25 Feb 2026, Dao et al., 13 Mar 2026) |
| Discrete-variable photonics | OAM quantum walks, frequency-bin electro-optic reservoirs, multimode fiber with photon coincidences | Single-setting informationally complete measurements, classical training with quantum inference, and interference-based feature maps (Zia et al., 25 Feb 2025, Brusaschi et al., 20 Mar 2026, Joly et al., 16 May 2025) |
| Continuous-variable photonics and GBS | Gaussian photonic QELM, GBS-based QELM | Fixed-time Gaussian substrates and high-dimensional sampling-based feature maps (Maier et al., 15 Oct 2025, Montesinos et al., 13 Jun 2026) |
In molecular modeling, QELM circuits were designed without extra ancilla qubits, using native single- and two-qubit gates, with qubit counts and circuit depths selected to match NISQ capabilities: 4–5 qubits and depth 26–27 for LiH, 5–7 qubits and depth 39–41 for water, and 7–9 qubits and depth 71–73 for formamide (Monaco et al., 2024). In exoplanet retrieval, each spectral patch is processed by a separate 5-qubit reservoir, parallelized across 9 independent groups of 5 qubits on IBM Fez, with 20,000 shots per example and full-dataset processing in about 6 hours (Vetrano et al., 3 Sep 2025). In digital superconducting QELM at utility scale, kicked-Ising circuits reached 124 qubits, 91 two-qubit-depth layers, and 5,084 two-qubit gates on ibm_quebec (Dao et al., 13 Mar 2026).
Photonic QELMs exhibit particularly diverse encodings and readouts. A photonic entanglement-witnessing QELM uses orbital angular momentum as an ancillary degree of freedom to realize informationally complete single-setting measurements of polarization entanglement, without requiring fine-tuning, precise calibration, or refined knowledge of the apparatus (Zia et al., 25 Feb 2025). A distinct frequency-bin implementation trains the QELM exclusively with intense classical fields through the correspondence between stimulated and spontaneous emission, and then performs inference on previously unseen quantum input states (Brusaschi et al., 20 Mar 2026). Another photonic realization uses indistinguishable photon pairs and a multimode fiber as a random densely connected layer, with output features taken from coincidence measurements (Joly et al., 16 May 2025). Continuous-variable photonic QELMs instead encode inputs through quadrature displacements and use Gaussian-compatible measurements to generate a high-dimensional random feature map at fixed latency (Maier et al., 15 Oct 2025).
5. Scientific and engineering applications
QELMs have been used extensively for quantum property learning. In photonic entanglement witnessing, a QELM trained on measurement statistics estimates entanglement witnesses directly and automatically adapts to noise and imprecisions while avoiding overfitting (Zia et al., 25 Feb 2025). In Werner-state estimation, a sequence of random Werner states is evolved with a reservoir state under an Ising Hamiltonian, local 7 observables form the features, and the protocol estimates the mixing parameter 8, which determines whether the state is entangled 9 (Assil et al., 3 Nov 2025). The same study reports that the optimal performance is found near the critical region rather than deep in the ergodic regime, and that adding two-point correlations further reduces MSE under noise (Assil et al., 3 Nov 2025). In a frequency-bin photonic QELM trained classically and evaluated quantumly, entanglement witnessing of two-qubit states reached 0 accuracy, multi-dimensional entanglement detection was demonstrated, and Hamiltonian learning reached fidelity 1 (Brusaschi et al., 20 Mar 2026).
Open-system characterization is another major application. QELMs have been used for quantum channel discrimination between Markovian and non-Markovian dynamics and for estimation of coupling strength 2 and depolarization rate 3 in tunable collision models (Assil et al., 17 Mar 2026). The constructive role of temporal memory is especially pronounced at low 4, where the dynamics are strongly non-Markovian (Assil et al., 17 Mar 2026). More generally, state-estimation QELMs remain effective beyond the scrambling time, and long-time performance becomes insensitive to the fine structure of the reservoir topology, matching Haar-random global unitaries across the tested settings (Vetrano et al., 2024).
In molecular science, QELMs have been applied to learn potential energy surfaces and force fields for LiH, water, and formamide, with all training restricted to a classical least-squares solve (Monaco et al., 2024). Under ideal statevector simulation, the reported RMSE(E) values are 5 for LiH, 6 for water, and 7 for formamide; on IBM Brisbane, the corresponding values are 8, 9, and 0 (Monaco et al., 2024). The same work reports lower RMSE and much lower circuit depth than VQE-based baselines for LiH and 1 (Monaco et al., 2024).
In astrophysics, a QELM framework for exoplanetary atmosphere retrieval uses spectral patching, PCA compression, and shallow parallel reservoirs on IBM Fez (Vetrano et al., 3 Sep 2025). On hardware, the reported retrieval accuracies are 100% for radius, 86% for 2, 84% for 3, 75% for 4, 40% for CO, 52% for mass, and 65% for temperature (Vetrano et al., 3 Sep 2025). The paper attributes the observed resilience to the shallow, parallel architecture and small-register design (Vetrano et al., 3 Sep 2025).
Industrial and data-intensive classical tasks have also been targeted. In the QUELL framework for elevator software, QELM is used for waiting-time prediction and statistically outperforms SVM and regression-tree baselines while maintaining performance under strong feature reduction (Wang et al., 2024). In a broader software-testing study covering Orona, Karie, and CaReSS, ideal-noiseless QELMs can outperform or match classical baselines, but noise on current hardware or noise models causes large degradations: more than 250% MSE increase in the regression task and more than 50% accuracy drop in the classification tasks when noise is applied at test time only (Muqeet et al., 2024). In collider-data selection, continuous-variable photonic QELMs outperform a parameter-matched MLP with two hidden units for all considered training sizes and match or exceed an MLP with ten hidden units at large sample sizes, while training only the linear readout (Maier et al., 15 Oct 2025). In GBS-based QELMs, photon-number sampling probabilities provide the best performance among the tested feature families, and on MNIST a 12-mode setup reached 96.9% median test accuracy using 2- and 4-photon probabilities (Montesinos et al., 13 Jun 2026). In multimode-fiber photonic QELM, coincidence-based protocols outperform intensity-based protocols in binary image classification, and simulations indicate that increasing the number of photons reveals a clear quantum advantage associated with increased rank of the feature matrix (Joly et al., 16 May 2025).
6. Limitations, concentration barriers, noise, and design criteria
A large part of the QELM literature is devoted to limitations. One limitation is classical simulability. For XX-Hamiltonian QELMs, the classification accuracy shows a relatively sharp transition from a low-accuracy to a high-accuracy regime at a critical time 5, after which accuracy saturates; the saturation value matches that achieved with random unitaries; and 6 is independent of the number of qubits (Lorenzis et al., 8 Sep 2025). Because information need only reach nearest neighbors, shallow local circuits and Matrix Product States can efficiently simulate the relevant regime, implying that no quantum computational advantage is expected for those tasks unless the task or the measurement becomes fundamentally more challenging (Lorenzis et al., 8 Sep 2025).
Another limitation is concentration. Four sources are identified that can cause exponential concentration of observables as system size grows: randomness, hardware noise, entanglement, and global measurements (Xiong et al., 2023). Under these conditions, observables can become input-agnostic, and with polynomially many shots it becomes statistically impossible to distinguish outputs for different inputs (Xiong et al., 2023). The same analysis argues against highly random reservoirs drawn from 2-design-like ensembles, and recommends intermediate chaoticity, local measurements, shallow encodings, and enough independent observables to avoid concentration while preserving expressivity (Xiong et al., 2023).
Noise and conditioning are persistent operational constraints. Finite measurement statistics induce estimation errors, and the amplification of those errors is governed by the condition number of the regression problem (Innocenti et al., 2022). In software-testing case studies, QELMs are significantly affected by quantum noise; training and testing with noise can partially reduce the gap, and error mitigation can lower the average performance drop in classification to 3.0%, but the effectiveness remains context dependent and insufficiently general for practical deployment across tasks (Muqeet et al., 2024). These results reinforce the distinction between ideal-simulator performance and hardware performance in the NISQ regime.
Recent work has responded to these barriers with more explicit design strategies. Large-scale digital QELM experiments on IBM Quantum processors introduce a multi-objective hyperparameter tuning strategy that jointly monitors observable variability, capacity, and task performance, together with a local eigentask analysis for scalable feature selection (Dao et al., 13 Mar 2026). This approach reports evidence of a regime of optimality identifiable at small scales and transferable across tasks and larger systems, with performance competitive with leading classical baselines on time-series forecasting and satellite image classification (Dao et al., 13 Mar 2026). A plausible implication is that practical QELM design is shifting from ad hoc reservoir choice toward principled co-design of encoding, dynamical regime, measurement locality, and readout robustness.
Across the literature, several design principles recur. Encoding determines the available nonlinear features; measurements determine which of those features are accessible; highly random or global strategies can induce concentration; and temporal or distributed augmentation can add useful structure without requiring fully variational training (Xiong et al., 2023, Gross et al., 20 Feb 2026, Gili et al., 12 Feb 2026). The field therefore presents QELM not as a single architecture, but as a family of fixed-quantum-feature-map models whose utility depends on how effectively a given physical substrate exposes task-relevant observables to a simple linear readout.