---
title: Cluster-Aware Upcycling in MoE Models
url: https://www.emergentmind.com/topics/cluster-aware-upcycling-91bf9079-c4ad-4c23-8129-71607be6f3d2
type: topic
---

# Cluster-Aware Upcycling in MoE Models

Cluster-aware Upcycling refers to a family of initialization and specialization techniques for Mixture-of-Experts (MoE) neural architectures, where pre-trained dense models are converted into sparse MoE structures with expert modules specialized according to the clustering structure of data or activations. This approach seeks to address expert symmetry, accelerate specialization, and improve transfer efficiency by aligning expert initialization and router dynamics with semantic or domain-specific clusters present in the underlying data. Cluster-aware upcycling has been investigated in both vision and multilingual language modeling, with distinct methodologies but common principles of leveraging cluster structure for expert allocation, initialization, and training stabilization [2604.13508][2604.25578].

## 1. Methodological Frameworks for Cluster-Aware Upcycling

Two primary instantiations of cluster-aware upcycling have been proposed: 
- Activation-based cluster upcycling for vision transformers [2604.13508], and 
- Language cluster–aware upcycling for multilingual MoEs [2604.25578].

### Activation-Based Semantic Clustering and Expert Initialization

Cluster-aware upcycling for vision models begins with collecting $\ell_2$-normalized activation vectors $X = \{x_j\}_{j=1}^M$ from a pre-trained feedforward network (FFN) block on a calibration corpus. Spherical $k$-means partitions $X$ into $K = N_e$ clusters by solving
$$
\{c_k\}_{k=1}^K = \arg\max_{\substack{\mu_1,\dots,\mu_K\\|\mu_i\|_2=1}} \sum_{j=1}^M \max_{1\le i\le K} \mu_i^T x_j,
$$
yielding semantic centroids $\{c_k\}$ and cluster blocks $X_i = \{x_j: i_j = i\}$. For each cluster, expert initialization proceeds via truncated SVD of a whitened projection $W S_i$, ensuring each expert $W_i$ best reconstructs the subspace of its cluster.

### Dense-to-MoE Upcycling with Fine-Grained Slicing and Drop-Upcycling

In large-scale multilingual MoEs, pre-trained dense FFN weights are partitioned into sub-matrices. For each layer, the weight matrices $W_{up}, W_{gate}, W_{down}$ are split along the hidden axis into $N$ slices, each forming the initialization for one expert. A scalar rescaling $\lambda = N^{1/3}$ preserves the expected forward-pass magnitude. Router weights are zero-initialized. To break initialization symmetry, "Drop-Upcycling" zeroes out random rows/columns in each expert's parameters and reinitializes them with noise matching the original mean and variance [2604.25578].

## 2. Cluster-Aware Routing and Expert Specialization Strategies

### Semantic Router Initialization

For vision MoEs, the router parameters $W_r$ are initialized row-wise with the $\ell_2$-normalized cluster centroids:
$$
W_r = \begin{bmatrix}
c_1^T \\ c_2^T \\ \vdots \\ c_{N_e}^T
\end{bmatrix} \in \mathbb{R}^{N_e \times d}.
$$
This ensures that, initially, inputs close to centroid $c_i$ are routed preferentially to expert $i$, catalyzing semantic alignment between router and expert.

### Data-Driven Expert Allocation and Emergent Clusters

In Marco-MoE, cluster-aware routing emerges from data-driven language-expert co-activation. Given a language $L$, expert activation profiles are recorded (number of tokens routed to expert $E_{i,\ell}$ for each layer). Pearson correlations of these activation vectors between language pairs define a language clustering metric $d(L_i, L_j) = 1 - \rho(L_i, L_j)$. Hierarchical clustering on these metrics empirically recovers linguistic families and reveals that related languages exhibit similar expert utilization [2604.25578].

## 3. Training Protocols and Cluster-Aware Losses

### Expert-Ensemble Self-Distillation

Cluster-aware upcycling for vision models introduces an expert-ensemble self-distillation (EESD) loss. An exponential moving average (EMA) teacher of all MoE parameters produces a dense ensemble output
$$
y_{ens}(x) = \sum_{i=1}^{N_e} g_i^{ema}(x) E_i^{ema}(x),
$$
against which the student's sparse output
$$
y_{MoE}(x) = \sum_{i \in \mathcal{T}_k(x)} \frac{g_i(x)}{\sum_{j \in \mathcal{T}_k(x)} g_j(x)} E_i(x)
$$
is regularized using either KL divergence on softmax predictions or squared error on hidden outputs:
$$
\mathcal{L}_{EESD} = \frac{1}{T} \sum_x \|\mathrm{sg}(y_{ens}(x)) - y_{MoE}(x)\|_2^2.
$$

### Router Regularization and Load-Balancing

In Marco-MoE, router training employs an auxiliary load-balancing loss ($\mathcal{L}_{balance}$) and a router Z-loss ($\mathcal{L}_z$), with coefficients set to $0.01$ and $0.001$, respectively. This prevents expert collapse and encourages uniform expert utilization.

## 4. Empirical Evaluation and Comparative Performance

### Quantitative Outcomes

Cluster-aware upcycling for vision transformers (CLIP ViT-B/32 and ViT-B/16 with 8-expert top-2 MoEs) consistently outperforms sparse upcycling baselines:

| Model/Metric         | Sparse Upcycling | Cluster-aware Upcycling | Gain           |
|:---------------------|:----------------:|:----------------------:|:--------------:|
| ViT-B/32 Retrieval   | (30.8, 48.0, 39.4) | (31.0, 48.2, 39.6)    | +0.2 avg.      |
| ViT-B/32 ImageNet-1k | 63.0%           | 63.2%                  | +0.2%          |
| ViT-B/16 Retrieval   | (34.9, 50.9, 42.9)| (35.4, 51.6, 43.5)    | +0.6 avg.      |
| VTAB-Natural         | 58.0%           | 59.1%                  | +1.1%          |

Few-shot and full fine-tuning further confirm consistent gains (e.g., 5-shot: 50.9 → 51.5, +0.6). In Marco-MoE, Nano and Mini models (0.6B–0.86B active params) match or exceed models with $1.7$–$4$B active parameters in both English and multilingual benchmarks, attaining best-in-class performance-to-compute ratios and competitive cultural/regional benchmark scores [2604.13508][2604.25578].

### Analysis of Expert Diversity and Routing

Cluster-aware initialization reduces inter-expert cosine similarity by $0.1$–$0.2$, lowers relative compactness (trace of within- vs. between-expert covariance) throughout training, and yields more orthogonal expert subspaces. Routing entropy per token is $\sim 20\%$ lower than baseline, reflecting greater confidence in assignment, yet expert utilization remains balanced and robust against collapse [2604.13508].

## 5. Clustering Dynamics, Scaling, and Domain Transfer

Language clustering analysis in Marco-MoE demonstrates that expert utilization signatures naturally encode phylogenetic language relationships. For languages within a detected cluster $C$, the most used experts (peak $p_C(i)$) are shared, while typologically distant languages activate distinct expert subsets. Hierarchical clustering accurately reconstructs known language family trees without explicit supervision. When expanding from 29 to 64 languages, expert activation signatures for pre-existing languages remain stable—indicating immunity to catastrophic interference during multilingual scaling [2604.25578].

A plausible implication is that cluster-aware upcycling methodologies can generalize to additional domains (e.g., programming languages, visual modalities), enabling new forms of structured specialization as long as domain clusters are present in the data.

## 6. Limitations, Open Questions, and Future Perspectives

Current cluster-aware upcycling approaches require full-model retraining for new domain or language expansion. Efficient modular expert addition without retraining, as well as improved router architectures (e.g., hierarchical or adaptive $k$) for sharper separation—especially for low-resource or micro-clusters—remain open challenges. Global scaling is limited by data availability for many low-resource languages or domains, suggesting that robust zero-shot adaptation or alternative specialization mechanisms will be necessary for comprehensive coverage [2604.25578].

In summary, cluster-aware upcycling provides an efficient blueprint for MoE initialization and specialization by structurally aligning experts and routers with data clusters. This enables greater expert diversity, early specialization, robust transfer, and efficiency in both vision and massive-scale multilingual MoE architectures [2604.13508][2604.25578].

Source: https://www.emergentmind.com/topics/cluster-aware-upcycling-91bf9079-c4ad-4c23-8129-71607be6f3d2