Loss of Trainability (LoT) Mechanisms
- Loss of Trainability (LoT) is a regime where optimization updates cease to yield significant improvements despite sufficient model capacity, expressivity, or supervision.
- LoT manifests through mechanisms such as exponentially suppressed gradients in quantum models, dead neurons in ReLU networks, and fractal boundaries in learning-rate space.
- Effective diagnosis of LoT requires a multi-metric approach that balances geometry, noise, and curvature rather than relying solely on single statistical proxies.
Searching arXiv for the cited LoT-related papers and closely related work to ground the article in current literature. Loss of Trainability (LoT) denotes the onset of regimes in which optimization updates cease to produce reliable improvement despite nominal model capacity, expressivity, or task supervision remaining adequate. Across the literature, the phrase refers not to a single mechanism but to a family of trainability failures whose operational signatures include exponentially suppressed gradients in variational quantum models, fractal trainability boundaries in learning-rate space for non-convex optimization, critical slowing in mechanically trained materials, dead-neuron–induced loss of effective capacity in ReLU networks, and information-flow collapse in deep feedforward or transformer architectures. Recent work emphasizes that LoT is architecture-dependent and often cannot be inferred from any single proxy such as expressivity, sharpness, Hessian rank, or weight norm alone (Mandal et al., 17 Jun 2026, Liu, 2024, Baveja et al., 24 Sep 2025).
1. Conceptual scope and formal criteria
In variational quantum algorithms (VQAs), trainability is usually defined through the behavior of the cost-function gradients for a parametrized quantum state
with cost
The relevant quantities are the partial derivatives , their variance over random initializations, and aggregate gradient norms such as . LoT appears when these gradients become too small to support efficient optimization, most prominently in barren plateau regimes where
or more generally in the number of qubits (Mandal et al., 17 Jun 2026). Closely related criteria are used in dissipative quantum neural networks, where trainability is likewise controlled by the scaling of gradient variance, and barren plateaus correspond to exponentially small variances such as or in the number of qubits (Sharma et al., 2020).
In quantum machine learning, the same barren plateau logic applies to losses built from datasets rather than single Hamiltonian objectives. For a generic loss
0
LoT is identified with exponential suppression of 1 or of related quantities such as Fisher-information entries, so that optimization requires exponential resources. The analysis in this setting shows that VQA trainability results extend to QML losses and that dataset structure can create additional trainability failures not present in standard VQAs (Thanasilp et al., 2021).
In classical deep learning, the notion is broader. One formulation treats trainability as the probability that a randomly initialized ReLU network contains sufficiently few permanently dead neurons to solve the task. If 2 denotes the number of permanently dead neurons at initialization, trainability is
3
for a task-dependent threshold 4; low 5 represents LoT because successful training is then impossible for most initializations (Shin et al., 2019). Another formulation, developed for deep feedforward networks, identifies trainability with the preservation of input information over depth and uses a reconstruction cutoff 6 as a proxy for correlation length. In that setting, LoT occurs when network depth exceeds the information-propagation scale, so that forward dynamics erase useful signal before it reaches the output (Thurn et al., 2024).
Optimization-centered analyses of continual learning define LoT as the regime in which gradient steps no longer yield improvement as tasks evolve, so accuracy stalls or degrades even though capacity and supervision are adequate. In that framework, LoT is not reliably predicted by any single statistic such as Hessian rank, sharpness level, weight norm, gradient norm, gradient-to-parameter ratio, or unit-sign entropy; instead it is governed jointly by gradient noise and curvature volatility (Baveja et al., 24 Sep 2025).
2. Barren plateaus, locality, and the quantum-variational picture
The most established formalization of LoT in quantum models is the barren plateau. Standard barren plateau results tie exponentially small gradient variance to sufficiently random or deep parametrized circuits that approximate unitary 2-designs. In that regime, global randomization of the state space induces concentration of measure, and local parameter updates produce exponentially weak effects on the objective (Mandal et al., 17 Jun 2026).
A central refinement is the distinction between local statistical complexity and global randomization. In structured finite-depth variational circuits studied on the one-dimensional cluster–Ising model and a generalized toric code Hamiltonian, statistical signatures commonly associated with randomness emerge at moderate circuit depth without any corresponding loss of trainability. The work analyzes three diagnostics of state complexity: Porter–Thomas statistics for measurement probabilities,
7
entanglement-spectrum adjacent-gap ratios
8
with benchmark means 9 for Poisson, 0 for GOE, and 1 for GUE, and the inverse participation ratio
2
Across both models, increasing depth drives these quantities toward Haar-random or random-matrix benchmarks while gradient variance shows no rapid exponential decrease over the available sizes 3, and VQE optimization remains effective (Mandal et al., 17 Jun 2026).
The explanation is explicitly local. For a local cost
4
and a bounded local-depth circuit, 5 depends only on Hamiltonian terms inside the causal light cone of the gate 6. Since that region has bounded size independent of total 7, fixed-depth local circuits do not realize the full concentration-of-measure mechanism responsible for 8 gradient scaling (Mandal et al., 17 Jun 2026). This suggests that random-looking local diagnostics are insufficient to diagnose LoT; what matters is the onset of global scrambling.
The same structural distinction appears in dissipative perceptron-based quantum neural networks. Deep global perceptrons that form unitary 2-designs produce barren plateaus both for global and, in some schemes, even local costs, with gradient-variance bounds such as
9
for random parameterized quantum circuits, and 0 under parameter-matrix multiplication updates (Sharma et al., 2020). By contrast, shallow local perceptrons with local costs can avoid barren plateaus, for example yielding
1
in a toy 2 construction, or polynomial lower bounds in a mapped hardware-efficient ansatz for 3 (Sharma et al., 2020). Here again, locality and local cost functions preserve trainability.
Quantum reinforcement learning shows an additional dependence on how measurement outcomes are grouped into actions. In PQC-based policies, trainability is controlled by the variance of the log-policy gradient 4. The analysis finds both barren plateaus and gradient explosion, with the qualitative regime determined by the type of basis-state partitioning and the mapping of partitions onto actions. A trainable window with polynomial measurement cost is guaranteed when a polynomial number of actions is encoded via contiguous-like partitioning of basis states (Sequeira et al., 2024). This situates LoT in quantum policy gradients within the same locality-versus-globality logic: action encoding can turn a nominally local observable into an effectively nonlocal one.
3. Expressivity, complexity, and when trainability separates from power
A recurring misconception is that higher expressivity or stronger statistical complexity automatically entails LoT. Several recent results reject that equivalence. For finite-depth local variational circuits, standard complexity diagnostics do not determine trainability (Mandal et al., 17 Jun 2026). For Instantaneous Quantum Polynomial circuit Born machines (IQP-QCBMs), the relevant structure is instead the interaction between generator sets, kernel spectra, and initialization (Shen et al., 11 Feb 2026).
In IQP-QCBMs trained with the MMD loss
5
the variance of the characteristic-function derivatives admits exact formulas under symmetric i.i.d. initialization. Under uniform initialization, if 6 denotes the anti-commuting generator set for frequency 7 and
8
is the critical rank, then
9
Thus LoT occurs frequency by frequency when 0, while low-weight frequencies with 1 remain trainable (Shen et al., 11 Feb 2026).
Kernel choice then becomes decisive. Flat or high-frequency spectra weight many frequencies with large 2, producing barren plateaus, whereas low-weight-biased kernels, such as Gaussian kernels with bandwidth 3, place substantial mass on the trainable low-Hamming-weight frequencies (Shen et al., 11 Feb 2026). Small-variance Gaussian initialization offers a second mitigation route: if 4, then
5
so gradients remain inverse-polynomially large for all frequencies that anticommute with at least one generator (Shen et al., 11 Feb 2026). A later analysis of Gaussian initialization in IQP-QCBMs derives complementary results using Stein’s lemma and Gaussian concentration, showing that LoT is controlled by the interaction of initialization variance 6, ansatz overlap structure 7, and kernel bandwidth. The same work gives a lower bound
8
and a concentration inequality
9
which make explicit when Gaussian schemes avoid or encourage exponential concentration (Luca, 8 Jun 2026).
Quantum walk optimization algorithms (QWOA) provide the opposite extreme: trainability is guaranteed precisely because expressivity is limited. For QWOA, the dynamical Lie algebra generated by the cost and mixing Hamiltonians has dimension bounded by
0
where 1 is the number of distinct cost values over feasible solutions; for 2 problems this is polynomial in input size (Bridi et al., 7 Aug 2025). Combining this with DLA-based variance formulas yields
3
so barren plateaus do not occur (Bridi et al., 7 Aug 2025). Yet the same work shows that for many optimization problems QWOA must be overparameterized to solve or approximate them, because its best-case performance remains Grover-like: 4 This suggests a different trainability limitation: avoiding barren plateaus does not remove complexity-theoretic performance bottlenecks or overparameterization-induced optimization difficulty (Bridi et al., 7 Aug 2025).
4. Classical deep networks: dead neurons, information-flow collapse, and fractal trainability boundaries
In ReLU networks, LoT is tied to neuron death and effective capacity loss. A neuron with preactivation 5 is dead when
6
and the paper distinguishes tentative death from permanent death. Permanent death corresponds to neurons whose output is constant on the domain and whose gradients with respect to their parameters vanish identically (Shin et al., 2019). If 7 is the activity indicator at initialization and
8
then trainability is the probability
9
This quantity is a necessary condition for successful training and an upper bound on the success probability over random initializations (Shin et al., 2019). Within this framework, over-parameterization is both necessary and sufficient for minimizing training loss because it makes 0 arbitrarily close to one for the task-dependent number 1 of required active neurons (Shin et al., 2019). A data-dependent initialization that places separating hyperplanes through the data cloud reduces the per-neuron dead probability and raises 2 accordingly (Shin et al., 2019).
A different classical LoT mechanism appears in hyperparameter space. For gradient descent
3
Liu studies bounded-versus-divergent behavior as a function of the learning rate 4 and shows that simple non-convex perturbations of a quadratic already produce fractal trainability boundaries (Liu, 2024). For the additive perturbation
5
the trainability boundary in learning-rate space is non-fractal while the loss remains convex, but becomes fractal as soon as non-convexity appears. The critical roughness
6
has threshold
7
which coincides with the point where the minimum second derivative reaches zero and convexity is lost (Liu, 2024). The paper interprets roughness as the factor controlling the gradient’s sensitivity to parameter changes, and the resulting fractal trainability boundary means that arbitrarily small changes in learning rate can flip training from convergent to divergent (Liu, 2024). This suggests a form of LoT rooted in dynamical-systems sensitivity rather than in vanishing gradients.
Forward information flow yields yet another classical view. For deep feedforward tanh networks with random initialization 8, 9, reconstruction cascades provide an empirical probe of whether input information survives to depth 0. A shallow auxiliary decoder 1 is trained so that
2
and cascaded reconstructions
3
are compared with the original input (Thurn et al., 2024). Relative entropy and differential reconstruction entropy quantify information loss; the depth 4 at which reconstruction entropy saturates acts as a correlation-length proxy. LoT occurs when 5, meaning that the network depth exceeds the information-propagation scale and forward dynamics erase useful signal (Thurn et al., 2024). This situates LoT within the ordered/chaotic phase diagram of random deep networks rather than within explicit gradient statistics.
A related but more detailed geometric theory is developed for deep randomly initialized transformers. Token representations are modeled as an interacting particle system whose geometry is summarized by the dot-product matrix 6. Starting from a permutation-symmetric simplex
7
the paper derives mean-field update equations for the pair 8 under layer norm, self-attention, MLP, and residual branches (Cowsik et al., 2024). Two Lyapunov exponents govern trainability: an angle exponent 9, which determines whether tokens collapse to a line (0) or repel into a regular simplex (1), and a gradient exponent 2, defined through
3
The paper shows that minimal test loss occurs when both exponents vanish,
4
and interprets LoT as a double failure of geometry and gradient propagation: either representation collapse or over-chaotic repulsion, combined with vanishing or exploding backpropagated gradients (Cowsik et al., 2024).
5. Physical learning systems and trainability transitions
Material training provides a non-neural, non-quantum example of LoT as a genuine critical phenomenon. In a 2D disordered spring network, training consists of cyclically driving one source site and 5 target sites so that target strains match prescribed in-phase or out-of-phase responses. Complexity is quantified by the fraction
6
where 7 is the number of nodes (Bhaumik et al., 2022). The training error decays as a power law in the number of cycles 8: 9 As 0 increases, 1 decreases continuously and vanishes at a critical threshold 2, defining the limit of trainable responses (Bhaumik et al., 2022). Near this point,
3
and the convergence time to a target error follows a Vogel–Fulcher–Tammann-like form
4
This is LoT as critical slowing down rather than as gradient collapse.
The mechanism is spectral. Training attempts to create one dominant soft mode corresponding to the desired source–target response, but increasing response complexity also generates a proliferation of spurious low-frequency modes. The density of states obeys
5
and the characteristic low-frequency scale drifts with training cycles as
6
Participation-ratio analysis shows that many of these low-frequency modes are localized around atypical local structures in which adjacent bonds nearly align (Bhaumik et al., 2022). Thus LoT in this system is caused by spectral crowding near zero modes and training-induced material degradation rather than by insufficient parameterization. This suggests a broader principle: learning can fail because optimization dynamics generate competing soft directions that overwhelm the target mode.
6. Optimization-level diagnostics and mitigation beyond single-metric explanations
In continual learning with Adam, LoT is analyzed as a mismatch between effective step size, gradient noise, and curvature volatility. A central result is that single indicators such as Hessian rank, sharpness level, weight norms, gradient norms, gradient-to-parameter ratios, and unit-sign entropy are not reliable predictors across architectures and regularizers (Baveja et al., 24 Sep 2025). Instead, the paper derives two complementary critical-step conditions.
The first is a batch-size-aware gradient-noise bound. Writing the minibatch gradient as 7 with per-sample variance 8, expected descent is controlled by the step size 9 relative to
00
up to a curvature factor. Large batch size increases this bound by averaging noise, whereas large per-sample variance shrinks it (Baveja et al., 24 Sep 2025).
The second is a curvature-volatility-controlled bound. With Adam preconditioning, the relevant curvature scale is the normalized sharpness
01
and over a window 02 the volatility is
03
where 04 and 05 are the mean and variance of 06 (Baveja et al., 24 Sep 2025). This yields a safe-step scale
07
Combining the two, the paper defines a per-layer volatility-inflated noise proxy
08
and critical effective step
09
A per-layer scheduler then cools or warms Adam’s base learning rate so that each layer’s effective step remains below a safe fraction of 10, stabilizing training across CReLU, Wasserstein regularization, and L2 weight decay (Baveja et al., 24 Sep 2025). This suggests that LoT is frequently a control problem: updates become unproductive not because any one landscape statistic is extreme, but because the optimizer’s effective step drifts outside a dynamically safe region.
An analogous control perspective appears in structured pruning. Here LoT arises because pruning damages dynamical isometry, making retraining strongly dependent on the learning rate. Trainability Preserving Pruning (TPP) addresses this by penalizing correlations between filters marked for pruning and retained filters using the Gram-matrix term
11
and by regularizing batch-normalization parameters of the pruned channels via
12
yielding total objective
13
with gradually increased 14 (Wang et al., 2022). On a 7-layer linear MLP, plain 15 pruning reduces the mean Jacobian singular value from 2.4987 in the unpruned network to 0.0040 after pruning, while TPP restores it to 3.4875 and matches the oracle trainability-recovery scheme in accuracy (Wang et al., 2022). This suggests that LoT after pruning is not purely a sparsity effect; it is mediated by how pruning alters inter-filter dependencies and normalization geometry.
A related optimization-theoretic result exists for attention and LoRA. For a single self-attention layer and for LoRA-parameterized shallow networks, adding arbitrarily mild super-quadratic regularization makes the empirical loss satisfy the Villani condition and induces a Poincaré inequality for the associated Gibbs measure
16
As a consequence, the Langevin-type SDE
17
converges in expectation to within 18 of the global minimum in 19 time for any data and architecture size (Sun et al., 8 May 2026). In the LoRA case, the unregularized factorization 20 has a non-compact scaling symmetry
21
which makes the Gibbs measure non-normalizable; mild factor regularization removes this degeneracy and restores trainability (Sun et al., 8 May 2026). This suggests that LoT can arise from global geometric pathologies of parameter space even when local gradients remain well defined.
7. Synthesis and recurring principles
The literature does not support a single universal mechanism for Loss of Trainability. Instead, LoT emerges whenever the mapping from parameter updates to objective improvement breaks down, and that breakdown can occur for structurally different reasons across domains.
One recurring principle is the distinction between local structure and global randomization. In finite-depth local VQAs, random-matrix-like entanglement spectra, Porter–Thomas statistics, or Haar-like inverse participation ratios can appear without barren plateaus because locality bounds the causal region of each gradient (Mandal et al., 17 Jun 2026). In IQP-QCBMs, similarly, trainability depends on generator topology and kernel spectrum, not on an undifferentiated notion of “quantum complexity” (Shen et al., 11 Feb 2026). This suggests that many statistical diagnostics are neither necessary nor sufficient indicators of LoT.
A second principle is that trainability is often governed by geometry and symmetry. ReLU networks lose trainability when too many units are permanently dead (Shin et al., 2019). Deep feedforward networks lose trainability when forward propagation destroys mutual information with the input before the last layer (Thurn et al., 2024). Transformers lose trainability when initialization pushes them away from the double-critical point 22, causing either line collapse or over-chaotic repulsion together with vanishing or exploding gradients (Cowsik et al., 2024). LoRA can lose trainability because of non-compact scaling orbits in factor space (Sun et al., 8 May 2026). These are all geometric rather than purely statistical failures.
A third principle is that LoT often reflects a mismatch between optimization scale and effective curvature/noise. Fractal trainability boundaries arise when non-convex roughness makes gradient descent hypersensitive to the learning rate (Liu, 2024). Continual-learning LoT with Adam is predicted by crossings of per-layer effective-step thresholds determined jointly by gradient noise and curvature volatility (Baveja et al., 24 Sep 2025). Pruning-induced LoT arises when the post-pruning network no longer supports stable gradient flow under the inherited retraining schedule (Wang et al., 2022). This suggests that mitigation often requires controlling step scales, curvature, or parameter-space symmetries rather than simply increasing model size.
Finally, trainability and capability are not monotone. QWOA avoids barren plateaus because its dynamical Lie algebra is too small to realize highly expressive random-like behavior, but this same limitation imposes Grover-style depth lower bounds and overparameterization for many hard problems (Bridi et al., 7 Aug 2025). Conversely, finite-depth local variational circuits can display strong local complexity while remaining trainable (Mandal et al., 17 Jun 2026). A plausible implication is that LoT is best understood not as a simple by-product of power or complexity, but as a failure of alignment between architecture, objective locality, initialization, and optimizer dynamics.