---
title: Parameter Group Freezing
url: https://www.emergentmind.com/topics/parameter-group-freezing
type: topic
---

# Parameter Group Freezing

Parameter group freezing is a structured approach for accelerating optimization, reducing memory/compute costs, or improving generalization by selectively fixing (i.e., freezing) disjoint subsets of model parameters during training or adaptation. The strategy is widely used in neural networks, quantum circuits, pipeline-parallel training, federated learning, continual learning, and lattice gauge simulations, with the specifics of partitioning, importance scoring, and scheduling varying by application domain.

## 1. Formal Definitions and Partitioning Strategies

Parameter group freezing entails partitioning the full parameter set $\theta$ into groups $\{\theta_g\}_{g=1}^{G}$, where each group represents either contiguous tensors (layers, adapters), functionally similar blocks (encoders, decoders), or arbitrary index subsets (coordinates, LoRA projections). For frozen groups, gradients are not computed or applied:
\[
\theta_g \leftarrow 
\begin{cases}
\theta_g - \eta \nabla_{\theta_g} \mathcal{L} & \text{if } g\ \text{is active} \\
\theta_g & \text{if } g\ \text{is frozen}
\end{cases}
\]
The identification and scheduling of frozen groups may occur statically (e.g., fixed decoder freezing in PEFT [2501.07818]) or dynamically, based on importance metrics computed during training (e.g., WSBD [2602.11383], AFLoRA [2403.13269], TimelyFreeze [2602.05754]).

Partitioning granularity is highly task-dependent:
- **Layer/group-wise:** e.g., entire Transformer decoders [2501.07818], RetinaNet heads [2402.12624], or pipeline-parallel stages [2602.05754].
- **Parameter-wise:** dynamic masking at the coordinate level (QNNs [2602.11383], FL [2504.01752]).
- **Module-wise LoRA projections:** separate freezing schedules for each low-rank adaptation matrix [2403.13269].

## 2. Importance Criteria and Group Selection

Critical to most adaptive freezing schemes is an “importance score” quantifying a parameter/group's contribution to learning/progress.

- **Gradient-accumulation metrics:** WSBD defines a per-parameter sum-of-gradients over a sliding window, $s_k = |\sum_{t=1}^\tau \partial \mathcal{C}/\partial \theta_k| + \epsilon$, assigning lowest-probability of activation to parameters with small magnitudes [2602.11383].
- **Activation statistics:** Layer mining for continual detection freezes groups according to (mean, median, variance, entropy) of activations [2402.12624].
- **Gradient-based stability:** AFLoRA computes EMAs of gradient magnitudes and their changes to generate a freezing score $s_\ell(t) = \text{mean}(m_\ell(t) \odot u_\ell(t))$, incrementally freezing LoRA projections with the smallest $s_\ell(t)$ [2403.13269].
- **Parameter drift:** Wireless FL monitors cumulative coordinate changes to determine which parameters have stabilized and are suitable for freezing [2504.01752].

The selection can be deterministic (freeze the lowest L% by some criterion) or stochastic (probabilistic masking as in WSBD). TimelyFreeze, in contrast, sets per-stage freeze ratios to optimize pipeline execution time under LP-imposed accuracy budgets [2602.05754].

## 3. Freezing Schedules and Adaptive Mechanisms

Several temporal strategies govern the dynamics of parameter group freezing:
- **Static freezing:** Certain groups remain frozen throughout (e.g., backbone layers in continual detection [2402.12624], frozen decoders in PEFT [2501.07818]).
- **Progressive freezing:** Groups become eligible for freezing as their importance scores indicate stabilization, employing staged or ramped schedules (AFLoRA's cubic ramp [2403.13269], TimelyFreeze's LP-based ramp [2602.05754]).
- **Windowed freezing/unfreezing:** Parameters are only allowed to update within fixed or adaptive windows, and reactivation is permitted to retain expressivity and avoid local minima (WSBD [2602.11383]).
- **Two-timescale frameworks:** In wireless FL, group freezing occurs at a large timescale (frames), while transmit power is dynamically controlled at finer granularity [2504.01752].

Pseudo-code for these routines typically involves: (1) computing importance metrics; (2) sorting/grouping; (3) applying mask/freeze status; (4) updating parameter subsets; and (optionally) (5) unfreezing/reactivating as scores change.

## 4. Theoretical Analysis and Convergence Guarantees

Freezing subgroups of parameters induces a bias-variance trade-off and raises questions around convergence and expressive capacity:

- **Bias from under-training:** Freezing introduces a persistent estimation bias proportional to the fraction and stability of frozen parameters. Federated learning bounds the total loss as $\mathbb{E}[F(w_T)] - F^* \leq A\prod_{t=1}^T(1-\mu\eta_t) + B\sum_{t=1}^T \eta_t^2(1+\rho y_t)$, where $y_t$ is the freeze fraction [2504.01752].
- **Hammerstein contraction in gradient masking:** Under L-smoothness and acceptance probability $p_{\min}$ for coordinate activation, masked-gradient SGD satisfies $p_\mathrm{min}\sum_{t}\eta_t \mathbb{E}[\|\nabla\mathcal{C}\|^2] < \infty$, and $\liminf_{t \to \infty} \mathbb{E}[\|\nabla\mathcal{C}\|^2]=0$ [2602.11383].
- **Pipeline efficiency:** TimelyFreeze frames the execution as an LP, constraining average per-stage freeze ratios to keep expected parameter updates above $1 - r_\mathrm{max}$, bounding iteration complexity via $T_\epsilon^{(\text{frozen})} \leq T_\epsilon^{(\text{base})}/(1 - r_\mathrm{max})$ [2602.05754].
- **Expressive capacity:** Unlike pruning, most freezing methods (except permanent parameter removal) preserve the possibility of full model adaptation, as frozen groups can be reactivated (WSBD) or limited to trainable adapters (AFLoRA).

## 5. Application Domains and Empirical Results

Parameter group freezing has been instantiated in a diverse set of practical machine learning and computational physics contexts:

| Domain                     | Grouping         | Key Outcomes               |
|----------------------------|------------------|----------------------------|
| Large Language Models      | Decoder layers   | ~50% fewer trainables, better multilingual retention, similar/better NLG metrics [2501.07818] |
| Continual Object Detection | Neck/head layers | Stability/plasticity tradeoff, best in class for mAP@.50 (e.g. 57.9 with 75% entropy-based freezing) [2402.12624] |
| Parameter-Efficient FT (PEFT) | LoRA projections | Up to $\mathbf{9.5\times}$ trainable parameter reduction, +0.85% average GLUE score, up to $1.86\times$ faster [2403.13269] |
| Quantum Neural Networks    | Parameters       | 63.9% faster than Adam, 80% forward-pass savings, robust to noise [2602.11383] |
| Wireless Federated Learning| Coordinates      | 15–20% total energy reduction for fixed accuracy, up to +3% accuracy for fixed energy [2504.01752] |
| Pipeline Parallelism       | Stage/microbatch | 36.3–40% throughput gain, $<$0.2% accuracy delta (LLaMA-8B) [2602.05754] |

For natural language generation, freezing decoders aligns with the pretraining distribution and avoids catastrophic forgetting, especially in multilingual regimes [2501.07818]. In object detection, hard layer freezing combined with entropy-based mining yields best performance/stability balance [2402.12624]. In quantum models, parameter-wise freezing via WSBD significantly mitigates the barren plateau effect [2602.11383]. In federated learning, coordinated two-timescale freezing and transmission control allow energy-bounded deployment with minimal loss in convergence [2504.01752].

## 6. Practical Recommendations and Limitations

Effective design and deployment of parameter group freezing require:

- **Appropriate importance metrics**: Select activation or gradient-based criteria suited to domain (e.g., mean activation for detection, sum-of-gradients for QNNs).
- **Careful freeze fraction selection**: Excessive freezing (e.g., $>70-80\%$) can severely degrade accuracy/plasticity [2402.12624, 2504.01752]; moderate ranges yield most gains.
- **Dynamic/interleaved schedules**: Enable adaptation to training signal shifts, especially in multitask and non-stationary learning [2602.11383, 2403.13269].
- **Hardware/process awareness**: Assess memory/compute tradeoffs (e.g., batch size limits, forward-pass counts).
- **Hyperparameter tuning**: Freeze window, warm-up period, penalty parameters, and LM thresholds can substantially affect performance.
- **Accuracy risks**: Freezing is generally safe if guided by reliable signals of parameter stabilization, but improper over-freezing or failing to adjust for uneven group sizes (e.g., neglecting Voronoi weights in discrete gauge simulations) leads to premature performance drop or even procedural biases [2501.07818, 2201.09625].

## 7. Specialized Variants: Freezing in Lattice Gauge Theory and Other Discrete Systems

Beyond machine learning, freezing transitions arise in digitized lattice gauge theory. There, constraints from finite group discretization (e.g., SU(2) subgroups) cause all but trivial configurations to be exponentially suppressed as coupling increases. The minimal angular separation $\Delta\theta$ among group elements controls the critical freezing threshold $\beta_\text{freeze}\sim\Delta\theta^{-2}$, requiring the use of "asymptotically dense" point sets (e.g., generalised Fibonacci spiral on $S^3$) and explicit weighting to avoid simulation freezing [2201.09625].

### Table: Freezing Transition Parameters in Discretized SU(2) Simulations

| Discretization     | Minimal Angle $\Delta\theta$ | $\beta_c$ at $n$ points         | Notes                                               |
|--------------------|------------------------------|----------------------------------|-----------------------------------------------------|
| Binary subgroups   | Fixed by group order         | $\sim N^2$ scaling               | Uniform weights, small $n$ ($n \leq 120$)           |
| Fibonacci lattice  | Decreases with $n$           | Matches refined Petcher-Weingarten| Near-optimal, arbitrary $n$, nearly uniform weights |

Practical guidance dictates choosing $n$ or $\Delta\theta$ to ensure $\beta_\text{max}(1 - \cos\Delta\theta_\text{min}) \lesssim O(1)$, and using generalised Fibonacci schemes for uniformity and maximal $\beta_c$ at a given computational cost [2201.09625].

## References

- "WSBD: Freezing-Based Optimizer for Quantum Neural Networks" [2602.11383]
- "A Two-Timescale Approach for Wireless Federated Learning with Parameter Freezing and Power Control" [2504.01752]
- "A Multi-Encoder Frozen-Decoder Approach for Fine-Tuning Large Language Models" [2501.07818]
- "Efficient Parameter Mining and Freezing for Continual Object Detection" [2402.12624]
- "AFLoRA: Adaptive Freezing of Low Rank Adaptation in Parameter Efficient Fine-Tuning of Large Models" [2403.13269]
- "TimelyFreeze: Adaptive Parameter Freezing Mechanism for Pipeline Parallelism" [2602.05754]
- "Digitising SU(2) Gauge Fields and the Freezing Transition" [2201.09625]

Source: https://www.emergentmind.com/topics/parameter-group-freezing