---
title: Selective Layer Training in Neural Networks
url: https://www.emergentmind.com/topics/selective-layer-training
type: topic
---

# Selective Layer Training in Neural Networks

Searching arXiv for recent and foundational papers on selective layer training and closely related layer-selective adaptation methods.
Selective layer training denotes a family of neural-network optimization and post-optimization strategies in which adaptation is not distributed uniformly across all parameters. Instead, only selected layers, blocks, channels, matrices, or supervision pathways are updated, restored, merged, or otherwise manipulated. In the literature, this family serves several distinct purposes: reducing compute and memory relative to full end-to-end optimization, improving generalization in low-data or heterogeneous settings, preserving pretrained knowledge during continued pre-training, and recovering specific capabilities such as reasoning or diversity through targeted interventions rather than whole-model retraining [2302.06354] [2408.15600] [2607.01232] [2602.06665].

## 1. Conceptual scope and selection units

The term covers more than one technical regime. In one regime, the training algorithm explicitly restricts gradient updates to a subset of pretrained layers while freezing the remainder. SubTuning formalizes this as selective layer finetuning: only a carefully chosen subset of layers is trained, with the rest kept at pretrained values [2302.06354]. In another regime, layers are not merely frozen or updated selectively, but decoupled and optimized sequentially. The sequential training algorithm for feedforward networks trains one hidden layer at a time, temporarily attaching an output node at each stage, then discarding the temporary output parameters before proceeding to the next layer [1905.07490].

A broader selective-training perspective also includes methods whose selection unit is smaller or structurally different from an entire layer. Label Selection Layer (LSL) selects whether a worker annotation contributes to the loss, so the selection target is the crowd label rather than the prediction output or the whole sample [2308.10396]. Selective Convolution reallocates parameter capacity among channels inside a convolutional layer through channel de-allocation and re-allocation, so the relevant unit is the channel rather than the full layer [1905.04509].

A further extension appears in training-free interventions performed after standard training has already finished. LASER replaces one selected matrix in one selected transformer layer with a low-rank approximation computed by SVD [2312.13558]. Selective Layer Restoration (SLR) restores a contiguous interval of layers in a post-trained LLM back to their pretrained weights [2602.06665]. MERIT selectively merges self-attention layers between a video-language model and its paired text-only backbone [2604.11399]. These methods are not fine-tuning in the ordinary sense, but they preserve the core selective-layer principle: model behavior is altered through targeted intervention on a chosen subset of the stack.

Taken together, the literature treats selective layer training less as a single algorithm than as an organizing idea. The central question is always which representational substructures should remain plastic, which should remain fixed, and which should be edited or reused.

## 2. Optimization formulations and theoretical principles

Several papers make explicit that selective layer training changes the optimization problem itself rather than merely reducing cost. In the sequential training algorithm, the full network objective
$$
\minimise_{V, W, X, y, \boldsymbol{\theta}, \boldsymbol{\phi}, \boldsymbol{\psi}, \omega} \sum_i \|z_i - z_i^*\|
$$
is replaced by a sequence of smaller subproblems, each optimizing one hidden layer plus a temporary output node. In the paper’s example, the full problem has \(37\) unknowns, whereas the first sequential stage has \(13\) and later stages have \(16\). Earlier temporary output parameters are “not used further but deserted,” and the final network is assembled by stacking the separately trained transformations [1905.07490].

SubTuning gives a formal sample-complexity argument for selective layer finetuning. Under a first-order Taylor approximation around pretrained parameters \(\theta\), the model becomes linearized in an NTK-style feature map, and the paper argues that full finetuning incurs generalization on the order of
$$
O\!\left(\frac{\sqrt{r}\,\Delta}{\sqrt{m}}\right),
$$
whereas tuning only a selected subset with parameter count \(r'\) yields
$$
O\!\left(\frac{\sqrt{r'}\,\Delta \log(kL)}{\sqrt{m}}\right).
$$
The logarithmic factor reflects the search over subsets. The formal implication is that restricting the trainable set can reduce effective complexity, especially when the number of training samples \(m\) is small [2302.06354].

The federated-learning analysis sharpens this point by showing that layer selection induces a biased optimization target unless all layers are selected identically by all clients. The paper defines a gradient mismatch
$$
\mathcal{E}_{t} \triangleq \left\| \nabla f(\theta^{t}) - \sum_{l\in \mathcal{L}_t} \nabla_l h_l^t(\theta^{t}) \right\|^2,
$$
and decomposes it into two terms: one from omitting important layers and one from heterogeneous client selections. Its convergence theorem implies an error floor of order \(\mathcal{O}(\mathcal{E}_{t,1} + \mathcal{E}_{t,2})\), so selective layer fine-tuning in FL may oscillate around a stationary point when important layers are omitted or client choices are poorly coordinated [2408.15600].

An analogous formalization appears in RL post-training for LLMs. The quantity layer contribution is defined as
$$
\mathcal{C}(k) = \frac{S_k - S_{\text{base}}}{S_{\text{full} } - S_{\text{base}}},
$$
measuring what fraction of the full-parameter RL gain is recovered by training only layer \(k\). This reframes layer selection as a measurable allocation problem: selective training is valuable only insofar as the chosen subspace captures the full-model improvement [2607.01232].

These formulations converge on a common principle. Selective layer training is not only a sparsity heuristic; it is an explicit restriction of the adaptation subspace, with corresponding effects on convergence, generalization, and transfer.

## 3. Selective fine-tuning and continued pre-training

SubTuning established an influential empirical pattern: finetuning different layers of a pretrained model can produce markedly different outcomes, and the best layer is not predictable simply from depth, parameter count, or spatial resolution. To exploit this, the paper defines a finetuning profile by tuning one layer or block at a time and then introduces Greedy SubTuning, which greedily adds the layer with the largest marginal validation gain. On VTAB-1k, selective tuning frequently outperforms full finetuning in the low-data regime. Representative results include ResNet-50 on CIFAR-100 at \(54.6\) for SubTuning versus \(33.7\) for full finetuning, and ViT-B/16 on Flowers102 at \(97.7\) versus \(91.2\). Under distribution shift on CIFAR-10-C with ResNet-26, SubTuning reports average accuracy \(84.2\), compared with \(81.1\) for full finetuning and \(82.0\) for surgical finetuning [2302.06354].

AdaBet addresses the same broad problem under a different constraint profile: on-device adaptation of pretrained vision models when labels, gradients, and memory are limited. It ranks layers using the normalized first Betti number
$$
\hat{b}_1^i = \frac{b_1^i}{|a^i|},
$$
computed from forward-pass activations alone. Layers are then selected according to a budget \(\rho\). On sixteen model–dataset pairs spanning ResNet50, VGG16, MobileNetV2, and ViT-B16 over Stanford Dogs, Oxford-IIIT Pets, CUB-200-2011, and Flowers102, AdaBet reports an average gain of \(5\%\) classification accuracy over the second-best baseline and an average \(40\%\) reduction in peak memory. The selection step is about \(45\%\) faster than ElasticTrainer for a single layer-selection step, and overall training is about \(11\%\) faster than full training on average [2510.03101].

AdaGradSelect studies selective block tuning for small language models. Its core score is the cumulative gradient magnitude within each transformer block,
$$
\text{importance}(b)=\sum_{w \in b}\| \nabla_w \mathcal{L} \|,
$$
combined with Dirichlet-based sampling and an \(\epsilon\)-greedy exploration schedule during the first epoch. In the reported SLM experiments on Qwen2.5-0.5B, LLaMA3.2-1B, and Phi4-mini-3.8B, the method trains about \(12\%\) faster, uses about \(35\%\) less GPU memory, and outperforms LoRA (\(r=256\)) on GSM8K by about \(3\%\) on average across models while remaining close to full fine-tuning accuracy [2512.15764].

Continued pre-training of LLMs introduces a different objective: preserve inherited knowledge while adapting cheaply to a new corpus. LayerTracer provides an interpretable diagnostic based on Task Particle and Layer-wise Sensitivity. It reports that deep layers act as critical regions for task execution and show higher stability, while shallow layers are more sensitive. The resulting recommendation is to train shallow layers and freeze deep layers. On Qwen3-710M-Base using the CCI3.0-HQ corpus, the train-shallow/freeze-deep strategy achieves \(29.40 / 29.55\) on C-Eval compared with \(26.62 / 26.32\) for full training, and \(28.25 / 28.42\) on CMMLU compared with \(26.62 / 26.32\) for full training [2605.11416].

Across these works, no single depth heuristic is universal. SubTuning rejects simple monotonic depth rules, AdaBet reports Betti peaks in early and late layers, and LayerTracer identifies deep layers as execution-critical but therefore better frozen than trained. This suggests that layer selection is regime-dependent: the best allocation depends on whether the goal is supervised adaptation, continued pre-training, or on-device retraining.

## 4. Federated and reinforcement-learning regimes

Selective layer fine-tuning has a distinct meaning in federated learning because different clients may select different layers under different budgets. The FL formulation assigns each client \(i\) a binary mask \(\mathbf{m}_i^t \in \{0,1\}^{L}\), with local updates
$$
\theta_i^{t,k} = \theta_i^{t,k-1} - \eta \sum_{l\in \mathcal{L}_i^t} g_{i,l}(\theta_i^{t,k-1}; \xi_{i}^{t,k-1}).
$$
The server then aggregates layer-dependent client updates. The proposed strategic selection method solves an objective that maximizes local gradient norms while penalizing disagreement in mask vectors across clients, subject to resource constraints. Empirically, under heterogeneous budgets \(R_i \in [1,4]\), the method is best across CIFAR-10, DomainNet, XGLUE-NC, and QA. Reported numbers include \(95.57\) versus \(95.43\) for Full on CIFAR-10 and \(65.80\) versus \(65.98\) for Full on QA, while preserving the flexibility of client-specific layer budgets [2408.15600].

In RL post-training for LLMs, the selective question shifts from client heterogeneity to layer contribution within a single transformer. The single-layer RL study evaluates seven models across Qwen3 and Qwen2.5 families, three RL algorithms, and task domains including mathematical reasoning, code generation, and agentic decision-making. Its central empirical result is that RL gains are highly concentrated in a small subset of layers, often in the middle of the stack. For Qwen3-1.7B, the best isolated layer is Layer \(10\) with \(\mathcal{C}=1.14\); for Qwen3-4B, Layer \(16\) reaches \(\mathcal{C}=1.06\); for Qwen3-8B, Layer \(16\) reaches \(\mathcal{C}=1.07\), while Layer \(0\) is negative at \(\mathcal{C}=-0.51\) [2607.01232].

The same work further reports that the middle-layer heuristic can outperform full-parameter RL when full profiling is unavailable. On Qwen3-8B, full RL achieves \(66.43 \pm 0.40\), whereas training only the top 10 empirically high-contribution layers reaches \(69.11 \pm 0.10\), and the middle-layer heuristic reaches \(68.19 \pm 0.62\). Layer rankings are also stable across tasks and datasets, with Spearman \(\rho=0.76\) between NuminaMath-CoT and DeepScaleR and \(\rho=0.59\) between NuminaMath-CoT and DeepCoder [2607.01232].

These two settings expose complementary constraints. In FL, selective tuning must respect cross-client aggregation geometry. In RL, the dominant issue is the nonuniform distribution of policy-improvement signal across the stack. In both cases, the literature rejects the assumption that uniform all-layer updating is the natural default.

## 5. Training-time selection below the whole-layer level

Selective training is also implemented below the level of whole-layer freezing. In learning from crowds, LSL inserts a label selector
$$
g(i,l,x,h),
$$
which determines whether a worker label contributes to the loss. The adoption rate is
$$
\phi(g|A)=\frac{1}{|A|}\sum_n\sum_{y_n^i\in A_n} g(i, y_n^i, x_n, h),
$$
and the selective empirical risk is
$$
\hat{r}(f, g|A)= \frac{ \frac{1}{|A|}\sum_n\sum_{y_n^i\in A_n} \ell(f(x_n), y_n^i)\cdot g(i, y_n^i, x_n, h) }{ \phi(g|A) }.
$$
The objective adds a coverage penalty
$$
\mathcal{L}(f, g)= \hat{r}(f, g|A) + \lambda \Psi(c - \phi(g|A)),
\qquad
\Psi(a) = (\max(0,a))^2.
$$
This mechanism is selective at the label/loss level rather than at inference time. Experimentally, the Feature-based Label Selection Layer performs better than Crowd Layer in all evaluation metrics on LabelMe, and all proposed LSL variants outperform any Crowd Layer variant on precision for CoNLL-2003 NER, while regression remains a limitation because LSL does not support scaled or shifted annotation results [2308.10396].

Selective Convolution modifies a standard convolutional layer into
$$
\mathrm{SelectConv}(\mathbf{X};\mathbf{W}) \coloneqq \mathrm{Conv}(\mathrm{SelectChannel}(\mathbf{X}); \mathbf{W}),
$$
where channel gates \(g_i\) can block channels and index variables \(\pi_i\) can duplicate important ones. To make duplication useful, the paper adds spatial shifting biases:
$$
\mathrm{SelectChannel}(\mathbf{X}; \mathbf{g}, \mathbf{b})_i \coloneqq g_i \cdot \mathrm{shift}(\mathbf{X}_{\pi_i}, b_i).
$$
Channel importance is estimated by the Expected Channel Damage Matrix (ECDM), and de-allocation is posed as a constrained optimization under an output-damage threshold \(\gamma\). The method dynamically redistributes capacity during SGD without increasing total parameter count. On CIFAR-10 with DenseNet-40, test error drops from \(6.62\%\) to \(6.09\%\); on ImageNet, DenseNet-121 improves from \(24.7\%\) to \(24.4\%\) and ResNet-50 from \(23.9\%\) to \(23.4\%\) [1905.04509].

These works show that selective training need not coincide with selective layer freezing. The selected object may be a crowd annotation, a channel slot, or a loss term. What unifies them is the allocation of optimization pressure toward parts of the training signal that are estimated to be most useful.

## 6. Training-free layer-selective modification after pretraining

A major branch of the literature dispenses with gradient-based retraining altogether. LASER performs a post-training SVD truncation of a chosen matrix type \(\tau \in \{q,k,v,o,in,out\}\) in a chosen layer \(\ell\), keeping only
$$
r = \lfloor \rho \cdot \mathrm{rank}_{\max} \rfloor
$$
singular components. The central empirical finding is that performance gains are concentrated in later MLP layers, especially the MLP input matrix. On CounterFact, GPT-J top-1 accuracy improves from \(13.3\%\) to \(24.1\%\) with a single-layer LASER intervention, and robustness to paraphrases increases by about \(24.8\) percentage points for questions already answered correctly. The improvement is not uniform across all metrics: on PILE, GPT-J MLP input perplexity changes from \(4.8\) to \(5.0\), indicating a small perplexity tradeoff [2312.13558].

Selective Layer Restoration addresses an opposite problem: post-training improves helpfulness and instruction-following but often causes mode collapse and diversity loss. SLR constructs a hybrid model \(M_{i:j}\) by restoring a contiguous interval of layers from the pretrained checkpoint:
$$
L_k = \begin{cases} 
L_k^{\text{pre}}, & i \le k \le j \\
L_k^{\text{post}}, & \text{otherwise}.
\end{cases}
$$
The interval is selected with the CRC proxy task under a quality floor \(Q_{\text{CRC}}(M_{i:j}) \ge 0.9 \cdot Q_{\text{CRC}}(M_{\text{post}})\). The chosen intervals are Llama \([12,17]\), Qwen \([12,27]\), and Gemma \([17,27]\). Across creative writing, SLR reports about \(+42.5\%\) diversity with about \(-2.1\%\) quality; on open-ended QA it reports about \(+112.7\%\) entropy and \(+169.5\%\) coverage-n with about \(-1.8\%\) precision; on reasoning it improves Pass@\(k\) at all \(k\) across Llama, Qwen, and Gemma [2602.06665].

MERIT extends layer-selective intervention to multimodal model merging. It selectively merges self-attention layers between a video-language model \(M\) and a paired text-only backbone \(N\), using a gate vector searched by CMA-ES and an objective
$$
\mathcal{F}(g) = \mathrm{Acc}_{\mathrm{TR}}(g) - \lambda \cdot D_{\mathrm{TP}}(g),
$$
which improves temporal reasoning while penalizing temporal-perception degradation. On Video-MME, LongVA-7B improves temporal reasoning from \(40.1\) to \(49.7\) while temporal perception improves from \(60.0\) to \(61.8\); InternVL3-8B improves temporal reasoning from \(52.0\) to \(57.6\) and temporal perception from \(74.5\) to \(81.8\) [2604.11399].

What these methods share is the claim that useful capability changes can be localized in layer space and recovered by editing the right subset rather than retraining the whole model. The practical consequence is that selective layer methods now encompass not only optimization-time allocation but also post hoc restoration, merging, and rank-reduction.

## 7. Limits, misconceptions, and unresolved issues

One common misconception is that the most important parameters are therefore the best ones to train. The Super Weights study directly rejects this inference. It shows that pruning certain scalar weights can be catastrophic, but training only those same coordinates is also disastrous. On ARC-Easy, training only top-\(k\) Super Weights in OLMo-1B or OLMo-7B, for \(k \in \{100,1000,4096,8192\}\), collapses accuracy to roughly random-guessing levels, about \(25\%\). Expanding to \(3 \times 3\) local neighborhoods up to \(36{,}864\) parameters does not help. By contrast, training \(4096\) random coordinates in the same down\_proj layers reaches \(64.18\%\) on OLMo-1B, above the \(60.65\%\) baseline, and vanilla LoRA reaches \(66.88\%\) on OLMo-1B and \(77.3\%\) on OLMo-7B with only \(0.16\%\) of parameters [2607.08733].

This result draws a sharp distinction between importance and trainability. The paper attributes the failure of Super Weight-only tuning both to the mismatch between sparse coordinate subspaces and the intrinsic low-dimensional fine-tuning subspace, and to the amplified curvature of Super Weight coordinates. Its conclusion is that effective PEFT relies on structured decompositions over entire layers rather than isolated high-importance coordinates [2607.08733].

Another misconception is that there is a universal answer to which layers should be trained. The surveyed literature does not support such a rule. SubTuning reports that the best layer is not predictable just from depth [2302.06354]. LayerTracer recommends training shallow layers and freezing deep layers in continued pre-training because deep layers are execution-critical and stable [2605.11416]. The RL study finds the highest-contribution layers in the middle of the stack [2607.01232]. MERIT’s best layer recipes are structured subsets that often span mid-to-late layers rather than a single contiguous “best region” [2604.11399]. A plausible implication is that selective layer training is governed by task-specific functional specialization rather than a fixed depth prior.

The literature also records explicit limitations. LSL is weaker than Crowd Layer in regression because it does not support scaled or shifted annotation results [2308.10396]. Selective Convolution relies on a Gaussian approximation of BN outputs and is most naturally specified for BN + ReLU CNN pipelines [1905.04509]. AdaBet is evaluated primarily on vision benchmarks and does not explicitly incorporate hardware-specific latency or energy profiles [2510.03101]. SLR experiments are limited to models around \(7\)–\(9\)B, and finer-grained restoration of heads or sub-layer blocks is left open [2602.06665]. In FL, heterogeneous client choices create a nonzero convergence floor unless carefully regulated [2408.15600].

The resulting picture is technically coherent but not reductive. Selective layer training is most effective when the trainable or editable degrees of freedom align with the geometry of adaptation, the deployment budget, and the functional role of the targeted layers. The field’s consistent finding is not that fewer parameters are always better, but that indiscriminate full-model updating is often unnecessary, and sometimes inferior, when layer-wise structure is used deliberately.

Source: https://www.emergentmind.com/topics/selective-layer-training