---
title: 'WSM: Model Merging in Weight Space'
url: https://www.emergentmind.com/topics/model-merging-wsm
type: topic
---

# WSM: Model Merging in Weight Space

Model merging in the weight space (abbreviated as WSM) denotes a family of techniques for integrating multiple fine-tuned or specialized neural network models into a single multitask model by direct manipulation and combination of their parameter vectors, without joint retraining or access to original training data. WSM methods have become pivotal in scenarios such as large language model (LLM) development, domain specialization, multi-task transfer, and federated learning, owing to their computational efficiency and ability to synthesize capabilities from heterogeneous sources.

## 1. Core Principles and Geometric Foundations

WSM is formalized over parameter sets $\{\theta^{(i)}\}_{i=1}^N$, all sharing a common architecture or a compatible alignment. The primary goal is to construct a merged parameter vector $\theta_{\mathrm{merged}}$ such that $f(x; \theta_{\mathrm{merged}})$ retains, and ideally extends, the predictive behaviors of the constituent models on their respective tasks. A central theme in modern WSM is the move from naive linear operations in parameter space to geometric or functional formulations that reflect distances in predictive distributions.

A key development is the formulation of model merging as minimization on the Fisher–Rao (FR) manifold, where the FR distance between models $\theta, \theta'$ is locally equivalent to the KL-divergence between their predictive distributions, i.e.,
$$
d^2_{\mathrm{FR}}(\theta, \theta') \approx (\theta-\theta')^T F(\theta) (\theta-\theta') \approx 2\,\mathrm{KL}(p_\theta||p_{\theta'}),
$$
where $F(\theta)$ is the Fisher information matrix. The optimal merge is a weighted Karcher (Fréchet) mean on the manifold:
$$
\theta^* = \arg\min_{\theta} \sum_{i=1}^{N} \alpha^{(i)} d_{\mathrm{FR}}(\theta, \theta^{(i)})^2
$$
subject to $\sum_i \alpha^{(i)} = 1$ and $\alpha^{(i)} \ge 0$ [2603.04972].

## 2. Major Algorithmic Paradigms

The landscape of WSM encompasses several algorithmic classes:

**a. Euclidean and Heuristic Approaches**  
Early WSM methods operate directly in Euclidean parameter space using linear averaging, task vector arithmetic, and uniform weighting. The general merge equation is:
$$
\theta_{\mathrm{merged}} = \sum_i \alpha_i \theta_i.
$$
Variants include per-layer or per-block weights, e.g., as optimized in evolutionary frameworks or via CMA-ES to maximize validation performance [2410.13699]. More advanced pruning-based heuristics (e.g., TIES, DARE) mask magnitude-insignificant or sign-conflicting parameters to mitigate destructive interference.

**b. Geometric and Information-Geometric Approaches**  
To overcome representation collapse with Euclidean methods, geometric approaches operate on the manifold defined by model functionals. The use of weighted Karcher means on the sphere (as a proxy for the Fisher-Rao manifold) ensures norm-preserving updates and prevents shrinkage of activation variance or effective matrix rank in deep layers, represented by blockwise updates on normalized parameter directions and re-scaling to source norms [2603.04972].

**c. Optimization and Data-Assisted Techniques**  
These approaches learn per-layer or per-parameter weights by minimizing multitask validation loss:
$$
\theta_{\mathrm{merged}}^{(j)} = \theta_{\mathrm{pre}}^{(j)} + \sum_{i=1}^k \tanh(w_{i,j})\,\Delta^{(i,j)},
$$
with gradient-based updates to $w_{i,j}$ [2412.10416], or parameter-wise interpolation coefficients $\alpha$ optimized in the presence of small supervised validation sets [2412.15467]. Convex quadratic programming (QP) over residual updates, using calibration data, yields globally optimal weights that minimize squared-output errors and generalize task arithmetic and soup approaches [2605.29101].

**d. Dynamic and Modular Recombination**  
Modern frameworks introduce modularization and dynamic input-aware routers. Components are decomposed into shared and task-exclusive modules, compressed (e.g., via SVD), and a lightweight router predicts the optimal weighting for each input, yielding input-conditioned composite models [2406.15479][2602.06552]. Component-wise and Pareto-efficient recombination balances storage and performance via multi-objective search.

**e. Bayesian and Covariance-Aware Methods**  
Bayesian model merging leverages anchor priors and activation-based Bayesian regression, with bi-level optimization (inner: module-wise closed-form MAP estimation; outer: global Bayesian optimization for hyperparameters) [2605.12843]. Data-free variants estimate required Gram matrices from weight differences, obviating the need for calibration data [2604.01329].

**f. Specialized and Heterogeneous Fusion**  
Scenarios requiring merging non-identical architectures utilize output distribution alignment, vocabulary mapping, and probabilistic fusion, allowing distributional behaviors to be merged without direct parameter combination [2410.13699]. Model Assembly Learning extends WSM to zookeeper-style layer-wise assembly with permutation-padded alignment between heterogeneous models [2503.21657].

The table below summarizes key classes:

| Paradigm                 | Methodological Principle      | Notable Example(s)           |
|--------------------------|------------------------------|------------------------------|
| Euclidean/Averaging      | Linear/interpolative in $\theta$ | Model Soup, Task Arithmetic    |
| Geometric/FR             | Karcher mean on FR manifold  | Spherical proxy fixed-point   |
| Optimization-based       | (Bi-level) supervised/validation tuning | SuperMerge, Output-Space QP  |
| Dynamic/Modular          | Component-wise, router-based | Twin-Merging, MERGE          |
| Bayesian                 | MAP with anchor priors       | Bayesian Model Merging        |
| Covariance-aware         | Layer-wise interference minimization | ACTMat                      |
| Distributional/heterogeneous | Output alignment/fusion | Unconstrained Model Merging   |

## 3. Empirical Performance and Scalability

State-of-the-art studies demonstrate several critical empirical properties:

- **Collapse Avoidance:** Geometry-aware or norm-preserving methods (e.g., Fisher–Rao Karcher mean) preserve activation variance and effective-rank, preventing the rapid degradation observed in naive Euclidean blends as the number and heterogeneity of experts increases. On Qwen2.5-14B, Karcher merging yields average task performance of 0.610 for $m=5$ experts vs. 0.542 (LERP) and 0.239 (Multi-SLERP) [2603.04972].
- **Combinatorial Reasoning:** Layer-wise or distribution-based fusion can produce emergent abilities, outperforming even the best individual source expert. For instance, merging math and code experts produces superior results on code-solving-math tasks, exceeding the performance of both original models [2410.13699].
- **Scalability Constraints:** Theoretical analysis predicts a sharp upper bound on the number of mergeable experts, with diminishing returns due to parameter space saturation (Gaussian width analysis) [2505.21226]. Empirically, expert count saturation typically occurs at 4–6 for vanilla merges, but can be extended with heavy-tailed parameter reparameterization and modularization.

## 4. Practical Algorithms, Implementation Guidelines, and Diagnostics

Efficient WSM necessitates careful trade-offs in computational, memory, and data requirements:

- **Block-wise/parallel merging** exploits tensor structure for both memory efficiency and parallelization. Merging is practical over $T\sim5$ iterations and $N\sim$ tens of models [2603.04972].
- **Memory management**: Hierarchical merging (e.g., tree-based merging of small subsets) bounds peak memory, enabling scaling to $k\sim 10$–$20$ [2412.10416].
- **Minimal validation requirements:** Many optimization-based methods require only tens of examples per-task for effective parameter estimation.
- **Performance diagnostics:** Variance, effective-rank preservation, and residual energy (fraction captured by output-space basis) act as forward diagnostics for merge reliability [2603.04972][2605.29101].
- **Tunable objectives:** Supervised merging weights, router architectures, SVD ranks, and merge duration (for checkpoint aggregation in pre-training) provide levers for balancing capacity, generalization, and compute.

## 5. Applications, Ecosystem, and Limitations

WSM is now widely used in:

- **Multi-task and instruction following LLMs:** Unifying domain-specific LLMs for math, code, translation, and safety alignment without retraining [2410.13699][2603.09938].
- **Federated and privacy-preserving learning:** Decentralized (disjoint, private) experts are merged without central data exposure [2412.15467].
- **Model compression and adapter fusion:** Low-rank LoRA modules or SVD-compressed experts enable memory-efficient storage and reversible merging [2510.14163].
- **Continual and decentralized LLM ecosystems:** Community-driven merging enables composition of public/private experts into next-generation models [2410.13699].
- **Dynamic routing and input-aware specialization:** Modular expert libraries with input-conditioned routers support efficient batch inference and on-demand adaptation [2406.15479][2602.06552].

Key toolkit and benchmark initiatives include MergeKit, FusionBench, and standardized multi-task evaluation suites.

Limitations and open challenges include:

- **Mergeability prediction:** Despite advances, a general theory predicting which fine-tuned models will merge successfully is incomplete. Empirical diagnostics such as base-model prior accuracy and empirical mergeability scores are currently the most reliable predictors [2601.06672].
- **Beyond architecture homogeneity:** Most methods assume identical, aligned architectures, although heterogeneous-architecture strategies exist via output distribution fusion [2410.13699] and padded/permuted assembly [2503.21657].
- **Scaling beyond $O(10)$ experts:** Theoretical and empirical saturation is a fundamental issue; only advanced approaches with dynamic, modular, or heavy-tailed augmentation scale further [2505.21226].
- **Cross-modal and generative scenario extension** remains a frontier.

## 6. Theoretical Bounds, Mergeability, and Failure Modes

Weight-space merging is inherently constrained by the dimensionality and effective geometry of the parameter space:

- **Mode connectivity and loss landscape geometry** are preconditions; interpolation is effective only when solutions reside within the same basin.
- **Mergeability** is highly non-uniform across fine-tuned updates; base model prior knowledge is the dominant predictor of a weight update’s survival probability under merging [2601.06672].
- **Upper bound on merged experts** is dictated by the residual variance: in uniform-merge scenarios with correlation $\rho$, the advantage over the base model decays as $O(1/n)$ in number of experts $n$.
- **Gaussian width analysis** shows marginal gain per added expert decays strictly concavely: fast initial benefit, followed by diminishing returns and statistical saturation [2505.21226].
- **Failure modes and mitigation:** Severe performance collapse arises from norm shrinkage, rank reduction, and parameter-space conflicts, especially with interfering, highly heterogeneous, or poorly aligned experts. Geometry-aware merging, weighting by inverse base-model accuracy, sparsification, and modularization can alleviate these effects.

## 7. Outlook and Research Frontiers

Model merging in weight space remains a fast-evolving research area. Near-term trajectories include the development of:

- **Dynamic, multi-granular, and adaptive merging strategies** integrating modular decomposition, input-aware routing, and context-dependent fusion.
- **Unified frameworks for cross-architecture, cross-modal, and multi-format expert merging**.
- **Predictive models of mergeability and automated optimization of merge strategies** leveraging geometric diagnostics and meta-learning.
- **Integration with safety, robustness, and continual adaptation protocols**, spanning both LLMs and multimodal systems.

WSM stands as a foundational tool for programmatically composing, extending, and distilling collective intelligence from distributed neural models across the modern AI landscape [2603.04972][2410.13699][2412.10416][2406.15479][2605.29101][2602.06552][2510.14163][2507.17634][2601.06672][2505.21226][2604.01329][2603.09938][2503.08998].

Source: https://www.emergentmind.com/topics/model-merging-wsm