Papers
Topics
Authors
Recent
Search
2000 character limit reached

Loss of Trainability (LoT) Mechanisms

Updated 12 July 2026
  • Loss of Trainability (LoT) is a regime where optimization updates cease to yield significant improvements despite sufficient model capacity, expressivity, or supervision.
  • LoT manifests through mechanisms such as exponentially suppressed gradients in quantum models, dead neurons in ReLU networks, and fractal boundaries in learning-rate space.
  • Effective diagnosis of LoT requires a multi-metric approach that balances geometry, noise, and curvature rather than relying solely on single statistical proxies.

Searching arXiv for the cited LoT-related papers and closely related work to ground the article in current literature. Loss of Trainability (LoT) denotes the onset of regimes in which optimization updates cease to produce reliable improvement despite nominal model capacity, expressivity, or task supervision remaining adequate. Across the literature, the phrase refers not to a single mechanism but to a family of trainability failures whose operational signatures include exponentially suppressed gradients in variational quantum models, fractal trainability boundaries in learning-rate space for non-convex optimization, critical slowing in mechanically trained materials, dead-neuron–induced loss of effective capacity in ReLU networks, and information-flow collapse in deep feedforward or transformer architectures. Recent work emphasizes that LoT is architecture-dependent and often cannot be inferred from any single proxy such as expressivity, sharpness, Hessian rank, or weight norm alone (Mandal et al., 17 Jun 2026, Liu, 2024, Baveja et al., 24 Sep 2025).

1. Conceptual scope and formal criteria

In variational quantum algorithms (VQAs), trainability is usually defined through the behavior of the cost-function gradients for a parametrized quantum state

ψ(θ)=U(θ)0,\ket{\psi(\boldsymbol{\theta})} = U(\boldsymbol{\theta})\ket{0},

with cost

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.

The relevant quantities are the partial derivatives θkE\partial_{\theta_k}E, their variance over random initializations, and aggregate gradient norms such as E2\langle |\nabla E|^2\rangle. LoT appears when these gradients become too small to support efficient optimization, most prominently in barren plateau regimes where

Var ⁣(Eθk)O(2N)\mathrm{Var}\!\left(\frac{\partial E}{\partial \theta_k}\right)\sim \mathcal{O}(2^{-N})

or more generally O(exp(cN))\mathcal{O}(\exp(-cN)) in the number of qubits NN (Mandal et al., 17 Jun 2026). Closely related criteria are used in dissipative quantum neural networks, where trainability is likewise controlled by the scaling of gradient variance, and barren plateaus correspond to exponentially small variances such as O(2n)\mathcal{O}(2^{-n}) or O(22n)\mathcal{O}(2^{-2n}) in the number of qubits nn (Sharma et al., 2020).

In quantum machine learning, the same barren plateau logic applies to losses built from datasets rather than single Hamiltonian objectives. For a generic loss

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.0

LoT is identified with exponential suppression of E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.1 or of related quantities such as Fisher-information entries, so that optimization requires exponential resources. The analysis in this setting shows that VQA trainability results extend to QML losses and that dataset structure can create additional trainability failures not present in standard VQAs (Thanasilp et al., 2021).

In classical deep learning, the notion is broader. One formulation treats trainability as the probability that a randomly initialized ReLU network contains sufficiently few permanently dead neurons to solve the task. If E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.2 denotes the number of permanently dead neurons at initialization, trainability is

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.3

for a task-dependent threshold E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.4; low E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.5 represents LoT because successful training is then impossible for most initializations (Shin et al., 2019). Another formulation, developed for deep feedforward networks, identifies trainability with the preservation of input information over depth and uses a reconstruction cutoff E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.6 as a proxy for correlation length. In that setting, LoT occurs when network depth exceeds the information-propagation scale, so that forward dynamics erase useful signal before it reaches the output (Thurn et al., 2024).

Optimization-centered analyses of continual learning define LoT as the regime in which gradient steps no longer yield improvement as tasks evolve, so accuracy stalls or degrades even though capacity and supervision are adequate. In that framework, LoT is not reliably predicted by any single statistic such as Hessian rank, sharpness level, weight norm, gradient norm, gradient-to-parameter ratio, or unit-sign entropy; instead it is governed jointly by gradient noise and curvature volatility (Baveja et al., 24 Sep 2025).

2. Barren plateaus, locality, and the quantum-variational picture

The most established formalization of LoT in quantum models is the barren plateau. Standard barren plateau results tie exponentially small gradient variance to sufficiently random or deep parametrized circuits that approximate unitary 2-designs. In that regime, global randomization of the state space induces concentration of measure, and local parameter updates produce exponentially weak effects on the objective (Mandal et al., 17 Jun 2026).

A central refinement is the distinction between local statistical complexity and global randomization. In structured finite-depth variational circuits studied on the one-dimensional cluster–Ising model and a generalized toric code Hamiltonian, statistical signatures commonly associated with randomness emerge at moderate circuit depth without any corresponding loss of trainability. The work analyzes three diagnostics of state complexity: Porter–Thomas statistics for measurement probabilities,

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.7

entanglement-spectrum adjacent-gap ratios

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.8

with benchmark means E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.9 for Poisson, θkE\partial_{\theta_k}E0 for GOE, and θkE\partial_{\theta_k}E1 for GUE, and the inverse participation ratio

θkE\partial_{\theta_k}E2

Across both models, increasing depth drives these quantities toward Haar-random or random-matrix benchmarks while gradient variance shows no rapid exponential decrease over the available sizes θkE\partial_{\theta_k}E3, and VQE optimization remains effective (Mandal et al., 17 Jun 2026).

The explanation is explicitly local. For a local cost

θkE\partial_{\theta_k}E4

and a bounded local-depth circuit, θkE\partial_{\theta_k}E5 depends only on Hamiltonian terms inside the causal light cone of the gate θkE\partial_{\theta_k}E6. Since that region has bounded size independent of total θkE\partial_{\theta_k}E7, fixed-depth local circuits do not realize the full concentration-of-measure mechanism responsible for θkE\partial_{\theta_k}E8 gradient scaling (Mandal et al., 17 Jun 2026). This suggests that random-looking local diagnostics are insufficient to diagnose LoT; what matters is the onset of global scrambling.

The same structural distinction appears in dissipative perceptron-based quantum neural networks. Deep global perceptrons that form unitary 2-designs produce barren plateaus both for global and, in some schemes, even local costs, with gradient-variance bounds such as

θkE\partial_{\theta_k}E9

for random parameterized quantum circuits, and E2\langle |\nabla E|^2\rangle0 under parameter-matrix multiplication updates (Sharma et al., 2020). By contrast, shallow local perceptrons with local costs can avoid barren plateaus, for example yielding

E2\langle |\nabla E|^2\rangle1

in a toy E2\langle |\nabla E|^2\rangle2 construction, or polynomial lower bounds in a mapped hardware-efficient ansatz for E2\langle |\nabla E|^2\rangle3 (Sharma et al., 2020). Here again, locality and local cost functions preserve trainability.

Quantum reinforcement learning shows an additional dependence on how measurement outcomes are grouped into actions. In PQC-based policies, trainability is controlled by the variance of the log-policy gradient E2\langle |\nabla E|^2\rangle4. The analysis finds both barren plateaus and gradient explosion, with the qualitative regime determined by the type of basis-state partitioning and the mapping of partitions onto actions. A trainable window with polynomial measurement cost is guaranteed when a polynomial number of actions is encoded via contiguous-like partitioning of basis states (Sequeira et al., 2024). This situates LoT in quantum policy gradients within the same locality-versus-globality logic: action encoding can turn a nominally local observable into an effectively nonlocal one.

3. Expressivity, complexity, and when trainability separates from power

A recurring misconception is that higher expressivity or stronger statistical complexity automatically entails LoT. Several recent results reject that equivalence. For finite-depth local variational circuits, standard complexity diagnostics do not determine trainability (Mandal et al., 17 Jun 2026). For Instantaneous Quantum Polynomial circuit Born machines (IQP-QCBMs), the relevant structure is instead the interaction between generator sets, kernel spectra, and initialization (Shen et al., 11 Feb 2026).

In IQP-QCBMs trained with the MMD loss

E2\langle |\nabla E|^2\rangle5

the variance of the characteristic-function derivatives admits exact formulas under symmetric i.i.d. initialization. Under uniform initialization, if E2\langle |\nabla E|^2\rangle6 denotes the anti-commuting generator set for frequency E2\langle |\nabla E|^2\rangle7 and

E2\langle |\nabla E|^2\rangle8

is the critical rank, then

E2\langle |\nabla E|^2\rangle9

Thus LoT occurs frequency by frequency when Var ⁣(Eθk)O(2N)\mathrm{Var}\!\left(\frac{\partial E}{\partial \theta_k}\right)\sim \mathcal{O}(2^{-N})0, while low-weight frequencies with Var ⁣(Eθk)O(2N)\mathrm{Var}\!\left(\frac{\partial E}{\partial \theta_k}\right)\sim \mathcal{O}(2^{-N})1 remain trainable (Shen et al., 11 Feb 2026).

Kernel choice then becomes decisive. Flat or high-frequency spectra weight many frequencies with large Var ⁣(Eθk)O(2N)\mathrm{Var}\!\left(\frac{\partial E}{\partial \theta_k}\right)\sim \mathcal{O}(2^{-N})2, producing barren plateaus, whereas low-weight-biased kernels, such as Gaussian kernels with bandwidth Var ⁣(Eθk)O(2N)\mathrm{Var}\!\left(\frac{\partial E}{\partial \theta_k}\right)\sim \mathcal{O}(2^{-N})3, place substantial mass on the trainable low-Hamming-weight frequencies (Shen et al., 11 Feb 2026). Small-variance Gaussian initialization offers a second mitigation route: if Var ⁣(Eθk)O(2N)\mathrm{Var}\!\left(\frac{\partial E}{\partial \theta_k}\right)\sim \mathcal{O}(2^{-N})4, then

Var ⁣(Eθk)O(2N)\mathrm{Var}\!\left(\frac{\partial E}{\partial \theta_k}\right)\sim \mathcal{O}(2^{-N})5

so gradients remain inverse-polynomially large for all frequencies that anticommute with at least one generator (Shen et al., 11 Feb 2026). A later analysis of Gaussian initialization in IQP-QCBMs derives complementary results using Stein’s lemma and Gaussian concentration, showing that LoT is controlled by the interaction of initialization variance Var ⁣(Eθk)O(2N)\mathrm{Var}\!\left(\frac{\partial E}{\partial \theta_k}\right)\sim \mathcal{O}(2^{-N})6, ansatz overlap structure Var ⁣(Eθk)O(2N)\mathrm{Var}\!\left(\frac{\partial E}{\partial \theta_k}\right)\sim \mathcal{O}(2^{-N})7, and kernel bandwidth. The same work gives a lower bound

Var ⁣(Eθk)O(2N)\mathrm{Var}\!\left(\frac{\partial E}{\partial \theta_k}\right)\sim \mathcal{O}(2^{-N})8

and a concentration inequality

Var ⁣(Eθk)O(2N)\mathrm{Var}\!\left(\frac{\partial E}{\partial \theta_k}\right)\sim \mathcal{O}(2^{-N})9

which make explicit when Gaussian schemes avoid or encourage exponential concentration (Luca, 8 Jun 2026).

Quantum walk optimization algorithms (QWOA) provide the opposite extreme: trainability is guaranteed precisely because expressivity is limited. For QWOA, the dynamical Lie algebra generated by the cost and mixing Hamiltonians has dimension bounded by

O(exp(cN))\mathcal{O}(\exp(-cN))0

where O(exp(cN))\mathcal{O}(\exp(-cN))1 is the number of distinct cost values over feasible solutions; for O(exp(cN))\mathcal{O}(\exp(-cN))2 problems this is polynomial in input size (Bridi et al., 7 Aug 2025). Combining this with DLA-based variance formulas yields

O(exp(cN))\mathcal{O}(\exp(-cN))3

so barren plateaus do not occur (Bridi et al., 7 Aug 2025). Yet the same work shows that for many optimization problems QWOA must be overparameterized to solve or approximate them, because its best-case performance remains Grover-like: O(exp(cN))\mathcal{O}(\exp(-cN))4 This suggests a different trainability limitation: avoiding barren plateaus does not remove complexity-theoretic performance bottlenecks or overparameterization-induced optimization difficulty (Bridi et al., 7 Aug 2025).

4. Classical deep networks: dead neurons, information-flow collapse, and fractal trainability boundaries

In ReLU networks, LoT is tied to neuron death and effective capacity loss. A neuron with preactivation O(exp(cN))\mathcal{O}(\exp(-cN))5 is dead when

O(exp(cN))\mathcal{O}(\exp(-cN))6

and the paper distinguishes tentative death from permanent death. Permanent death corresponds to neurons whose output is constant on the domain and whose gradients with respect to their parameters vanish identically (Shin et al., 2019). If O(exp(cN))\mathcal{O}(\exp(-cN))7 is the activity indicator at initialization and

O(exp(cN))\mathcal{O}(\exp(-cN))8

then trainability is the probability

O(exp(cN))\mathcal{O}(\exp(-cN))9

This quantity is a necessary condition for successful training and an upper bound on the success probability over random initializations (Shin et al., 2019). Within this framework, over-parameterization is both necessary and sufficient for minimizing training loss because it makes NN0 arbitrarily close to one for the task-dependent number NN1 of required active neurons (Shin et al., 2019). A data-dependent initialization that places separating hyperplanes through the data cloud reduces the per-neuron dead probability and raises NN2 accordingly (Shin et al., 2019).

A different classical LoT mechanism appears in hyperparameter space. For gradient descent

NN3

Liu studies bounded-versus-divergent behavior as a function of the learning rate NN4 and shows that simple non-convex perturbations of a quadratic already produce fractal trainability boundaries (Liu, 2024). For the additive perturbation

NN5

the trainability boundary in learning-rate space is non-fractal while the loss remains convex, but becomes fractal as soon as non-convexity appears. The critical roughness

NN6

has threshold

NN7

which coincides with the point where the minimum second derivative reaches zero and convexity is lost (Liu, 2024). The paper interprets roughness as the factor controlling the gradient’s sensitivity to parameter changes, and the resulting fractal trainability boundary means that arbitrarily small changes in learning rate can flip training from convergent to divergent (Liu, 2024). This suggests a form of LoT rooted in dynamical-systems sensitivity rather than in vanishing gradients.

Forward information flow yields yet another classical view. For deep feedforward tanh networks with random initialization NN8, NN9, reconstruction cascades provide an empirical probe of whether input information survives to depth O(2n)\mathcal{O}(2^{-n})0. A shallow auxiliary decoder O(2n)\mathcal{O}(2^{-n})1 is trained so that

O(2n)\mathcal{O}(2^{-n})2

and cascaded reconstructions

O(2n)\mathcal{O}(2^{-n})3

are compared with the original input (Thurn et al., 2024). Relative entropy and differential reconstruction entropy quantify information loss; the depth O(2n)\mathcal{O}(2^{-n})4 at which reconstruction entropy saturates acts as a correlation-length proxy. LoT occurs when O(2n)\mathcal{O}(2^{-n})5, meaning that the network depth exceeds the information-propagation scale and forward dynamics erase useful signal (Thurn et al., 2024). This situates LoT within the ordered/chaotic phase diagram of random deep networks rather than within explicit gradient statistics.

A related but more detailed geometric theory is developed for deep randomly initialized transformers. Token representations are modeled as an interacting particle system whose geometry is summarized by the dot-product matrix O(2n)\mathcal{O}(2^{-n})6. Starting from a permutation-symmetric simplex

O(2n)\mathcal{O}(2^{-n})7

the paper derives mean-field update equations for the pair O(2n)\mathcal{O}(2^{-n})8 under layer norm, self-attention, MLP, and residual branches (Cowsik et al., 2024). Two Lyapunov exponents govern trainability: an angle exponent O(2n)\mathcal{O}(2^{-n})9, which determines whether tokens collapse to a line (O(22n)\mathcal{O}(2^{-2n})0) or repel into a regular simplex (O(22n)\mathcal{O}(2^{-2n})1), and a gradient exponent O(22n)\mathcal{O}(2^{-2n})2, defined through

O(22n)\mathcal{O}(2^{-2n})3

The paper shows that minimal test loss occurs when both exponents vanish,

O(22n)\mathcal{O}(2^{-2n})4

and interprets LoT as a double failure of geometry and gradient propagation: either representation collapse or over-chaotic repulsion, combined with vanishing or exploding backpropagated gradients (Cowsik et al., 2024).

5. Physical learning systems and trainability transitions

Material training provides a non-neural, non-quantum example of LoT as a genuine critical phenomenon. In a 2D disordered spring network, training consists of cyclically driving one source site and O(22n)\mathcal{O}(2^{-2n})5 target sites so that target strains match prescribed in-phase or out-of-phase responses. Complexity is quantified by the fraction

O(22n)\mathcal{O}(2^{-2n})6

where O(22n)\mathcal{O}(2^{-2n})7 is the number of nodes (Bhaumik et al., 2022). The training error decays as a power law in the number of cycles O(22n)\mathcal{O}(2^{-2n})8: O(22n)\mathcal{O}(2^{-2n})9 As nn0 increases, nn1 decreases continuously and vanishes at a critical threshold nn2, defining the limit of trainable responses (Bhaumik et al., 2022). Near this point,

nn3

and the convergence time to a target error follows a Vogel–Fulcher–Tammann-like form

nn4

This is LoT as critical slowing down rather than as gradient collapse.

The mechanism is spectral. Training attempts to create one dominant soft mode corresponding to the desired source–target response, but increasing response complexity also generates a proliferation of spurious low-frequency modes. The density of states obeys

nn5

and the characteristic low-frequency scale drifts with training cycles as

nn6

Participation-ratio analysis shows that many of these low-frequency modes are localized around atypical local structures in which adjacent bonds nearly align (Bhaumik et al., 2022). Thus LoT in this system is caused by spectral crowding near zero modes and training-induced material degradation rather than by insufficient parameterization. This suggests a broader principle: learning can fail because optimization dynamics generate competing soft directions that overwhelm the target mode.

6. Optimization-level diagnostics and mitigation beyond single-metric explanations

In continual learning with Adam, LoT is analyzed as a mismatch between effective step size, gradient noise, and curvature volatility. A central result is that single indicators such as Hessian rank, sharpness level, weight norms, gradient norms, gradient-to-parameter ratios, and unit-sign entropy are not reliable predictors across architectures and regularizers (Baveja et al., 24 Sep 2025). Instead, the paper derives two complementary critical-step conditions.

The first is a batch-size-aware gradient-noise bound. Writing the minibatch gradient as nn7 with per-sample variance nn8, expected descent is controlled by the step size nn9 relative to

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.00

up to a curvature factor. Large batch size increases this bound by averaging noise, whereas large per-sample variance shrinks it (Baveja et al., 24 Sep 2025).

The second is a curvature-volatility-controlled bound. With Adam preconditioning, the relevant curvature scale is the normalized sharpness

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.01

and over a window E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.02 the volatility is

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.03

where E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.04 and E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.05 are the mean and variance of E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.06 (Baveja et al., 24 Sep 2025). This yields a safe-step scale

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.07

Combining the two, the paper defines a per-layer volatility-inflated noise proxy

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.08

and critical effective step

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.09

A per-layer scheduler then cools or warms Adam’s base learning rate so that each layer’s effective step remains below a safe fraction of E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.10, stabilizing training across CReLU, Wasserstein regularization, and L2 weight decay (Baveja et al., 24 Sep 2025). This suggests that LoT is frequently a control problem: updates become unproductive not because any one landscape statistic is extreme, but because the optimizer’s effective step drifts outside a dynamically safe region.

An analogous control perspective appears in structured pruning. Here LoT arises because pruning damages dynamical isometry, making retraining strongly dependent on the learning rate. Trainability Preserving Pruning (TPP) addresses this by penalizing correlations between filters marked for pruning and retained filters using the Gram-matrix term

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.11

and by regularizing batch-normalization parameters of the pruned channels via

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.12

yielding total objective

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.13

with gradually increased E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.14 (Wang et al., 2022). On a 7-layer linear MLP, plain E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.15 pruning reduces the mean Jacobian singular value from 2.4987 in the unpruned network to 0.0040 after pruning, while TPP restores it to 3.4875 and matches the oracle trainability-recovery scheme in accuracy (Wang et al., 2022). This suggests that LoT after pruning is not purely a sparsity effect; it is mediated by how pruning alters inter-filter dependencies and normalization geometry.

A related optimization-theoretic result exists for attention and LoRA. For a single self-attention layer and for LoRA-parameterized shallow networks, adding arbitrarily mild super-quadratic regularization makes the empirical loss satisfy the Villani condition and induces a Poincaré inequality for the associated Gibbs measure

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.16

As a consequence, the Langevin-type SDE

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.17

converges in expectation to within E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.18 of the global minimum in E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.19 time for any data and architecture size (Sun et al., 8 May 2026). In the LoRA case, the unregularized factorization E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.20 has a non-compact scaling symmetry

E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.21

which makes the Gibbs measure non-normalizable; mild factor regularization removes this degeneracy and restores trainability (Sun et al., 8 May 2026). This suggests that LoT can arise from global geometric pathologies of parameter space even when local gradients remain well defined.

7. Synthesis and recurring principles

The literature does not support a single universal mechanism for Loss of Trainability. Instead, LoT emerges whenever the mapping from parameter updates to objective improvement breaks down, and that breakdown can occur for structurally different reasons across domains.

One recurring principle is the distinction between local structure and global randomization. In finite-depth local VQAs, random-matrix-like entanglement spectra, Porter–Thomas statistics, or Haar-like inverse participation ratios can appear without barren plateaus because locality bounds the causal region of each gradient (Mandal et al., 17 Jun 2026). In IQP-QCBMs, similarly, trainability depends on generator topology and kernel spectrum, not on an undifferentiated notion of “quantum complexity” (Shen et al., 11 Feb 2026). This suggests that many statistical diagnostics are neither necessary nor sufficient indicators of LoT.

A second principle is that trainability is often governed by geometry and symmetry. ReLU networks lose trainability when too many units are permanently dead (Shin et al., 2019). Deep feedforward networks lose trainability when forward propagation destroys mutual information with the input before the last layer (Thurn et al., 2024). Transformers lose trainability when initialization pushes them away from the double-critical point E(θ)=ψ(θ)Hψ(θ).E(\boldsymbol{\theta}) = \bra{\psi(\boldsymbol{\theta})} H \ket{\psi(\boldsymbol{\theta})}.22, causing either line collapse or over-chaotic repulsion together with vanishing or exploding gradients (Cowsik et al., 2024). LoRA can lose trainability because of non-compact scaling orbits in factor space (Sun et al., 8 May 2026). These are all geometric rather than purely statistical failures.

A third principle is that LoT often reflects a mismatch between optimization scale and effective curvature/noise. Fractal trainability boundaries arise when non-convex roughness makes gradient descent hypersensitive to the learning rate (Liu, 2024). Continual-learning LoT with Adam is predicted by crossings of per-layer effective-step thresholds determined jointly by gradient noise and curvature volatility (Baveja et al., 24 Sep 2025). Pruning-induced LoT arises when the post-pruning network no longer supports stable gradient flow under the inherited retraining schedule (Wang et al., 2022). This suggests that mitigation often requires controlling step scales, curvature, or parameter-space symmetries rather than simply increasing model size.

Finally, trainability and capability are not monotone. QWOA avoids barren plateaus because its dynamical Lie algebra is too small to realize highly expressive random-like behavior, but this same limitation imposes Grover-style depth lower bounds and overparameterization for many hard problems (Bridi et al., 7 Aug 2025). Conversely, finite-depth local variational circuits can display strong local complexity while remaining trainable (Mandal et al., 17 Jun 2026). A plausible implication is that LoT is best understood not as a simple by-product of power or complexity, but as a failure of alignment between architecture, objective locality, initialization, and optimizer dynamics.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Loss of Trainability (LoT).