---
title: Interleaved Stacking in Speech Model Distillation
url: https://www.emergentmind.com/topics/interleaved-stacking
type: topic
---

# Interleaved Stacking in Speech Model Distillation

Interleaved stacking is a stagewise stacking method for speech foundation model distillation in which copied layers are inserted immediately after their source layers so that layer position is preserved as model depth increases. It was introduced to address a limitation of existing stacking methods: while shallow-to-deep growth reduces training cost, copying blocks to the top of the network repeatedly relocates layers and can degrade downstream performance in speech foundation models, where early, middle, and late layers encode distinct layer-specific knowledge. In this setting, interleaved stacking is not merely a schedule for reducing wall-clock time; it is a positional constraint on depth growth designed to preserve the semantic identity of layers during distillation [2606.11766].

## 1. Problem setting in speech foundation model distillation

The method was proposed in the context of distilling a large speech foundation model into an efficient student model for low-resource environments. Distillation reduces inference latency, but it requires additional student-model training, and the training efficiency of speech foundation model distillation had remained underexplored. The specific problem is therefore not whether distillation is useful, but how to accelerate student training without sacrificing the representational structure that makes distilled speech models effective [2606.11766].

The paper frames stacking as a natural candidate for training acceleration. If the target student has depth \(N\) and training is split into \(B\) stages, then a shallower model can be trained first and expanded progressively until the target depth is reached. Early stages are cheaper because they use fewer layers. However, in speech foundation models, this economy creates a structural risk: different layers specialize differently, and layer-specific knowledge matters. A stacking rule that changes the meaning of a layer’s depth position forces the model to relearn roles at each stage. This suggests that, in speech, speed-oriented depth growth must be constrained by representational continuity rather than treated as a purely architectural convenience [2606.11766].

## 2. Stagewise stacking and the problem of layer relocation

In standard stagewise stacking, \(K=N/B\), and at stage \(b\) the model has \(bK\) layers. The network grows from \(K\) to \(2K\) to \(3K\), and so on until \(N\). The efficiency benefit comes from training the shallower stages first. The difficulty lies in how new layers are initialized and placed [2606.11766].

Two prior strategies are described. Gradual stacking initializes the next stage by copying the last \(K\) layers and stacking them on top:
\[
f_{bK:1} \rightarrow f_{bK:(b-1)K+1}\circ f_{bK:(b-1)K+1}\circ f_{(b-1)K:1}.
\]
MIDAS instead copies an intermediate block:
\[
f_{bK:1} \rightarrow f_{bK:iK+1}\circ f_{iK:(i-1)K+1}\circ f_{iK:(i-1)K+1}\circ f_{(i-1)K:1},
\]
where \(i=\lceil b/2\rceil\) [2606.11766].

The common limitation is that copied layers are relocated into a new structural context. A layer that was late in one stage may become middle in the next stage. In a speech foundation model, that relocation is problematic because early and late layers do not behave as interchangeable “depth slots.” The paper therefore treats positional instability, rather than copying itself, as the main source of performance degradation [2606.11766].

| Method | Copy strategy | Positional effect |
|---|---|---|
| Gradual stacking | Copy the last \(K\) layers and stack them on top | Late layers can become middle layers |
| MIDAS | Copy an intermediate block | Copied block is relocated into a new context |
| Interleaved stacking | Copy evenly spaced layers and insert each after its source layer | Relative layer position is preserved |

A common misconception is that all stacking methods are equivalent up to training speed. The reported comparison rejects that view: the decisive variable is not only reduced computation in early stages, but whether depth growth preserves the representational role associated with layer position.

## 3. Interleaved stacking as positional preservation

Interleaved stacking differs from both gradual stacking and MIDAS by copying evenly spaced layers and inserting them immediately after their originals. In the \(b\)-th stage, it copies layer \(f_{bk}\) for \(k=1,\dots,K\), so that each copied layer is duplicated in place:
\[
f_{bk}=f_{bk}\circ f_{bk}.
\]
The paper emphasizes that the important point is not the algebraic identity by itself, but the placement rule: the copy stays adjacent to its source instead of being appended at the top [2606.11766].

Conceptually, this means that early, middle, and late layers continue to occupy roughly the same relative positions as the model expands. The copied layers are not “moved upward” to become a different part of the hierarchy. The method therefore preserves what the paper describes as layer identity across the entire depth-growth schedule. In speech foundation models, where layers encode different kinds of information at different depths, that consistency is presented as especially critical [2606.11766].

The central claim is thus narrower and more technical than a generic “better initialization” argument. Interleaved stacking is designed to preserve the mapping between depth position and function during stagewise expansion. A plausible implication is that the method is most relevant in settings where representational specialization across depth is strong, and less critical in architectures where layers are more nearly interchangeable.

## 4. Distillation objective and intermediate supervision

The method is embedded in a knowledge-distillation setup that combines output-level and intermediate-level supervision. The output-level KD loss is defined between the teacher output \(y^T\in\mathbb{R}^{F\times D^T}\) and the student output \(y\in\mathbb{R}^{F\times D}\), with a projection layer to match dimensions:
\[
\mathcal{L}=\mathrm{MSE}(y^T,\mathrm{proj}(y)).
\]
The total objective adds an intermediate-level KD term:
\[
\mathcal{L}_{\text{total}}=\mathcal{L}+w\,\mathcal{L}_{\text{inter}},
\]
with \(w\) controlling the strength of intermediate supervision [2606.11766].

For a \(K\)-layer initial-stage student, there are \(K-1\) intermediate losses, and the teacher targets are fixed to evenly spaced intermediate layers. The paper gives a 12-layer teacher/student example with \(B=4\), so \(K=3\): teacher layers 4 and 8 may be compared to student layers 1 and 2 in stage 1, then to student layers 2 and 4 in stage 2. Importantly, the teacher target indices are kept fixed across stages, which the paper presents as stabilizing and speeding up training [2606.11766].

This interaction between stacking and distillation is a central part of the proposal. Interleaved stacking works well with intermediate-level KD losses because student layers remain in consistent positions across stages. Gradual stacking, by contrast, changes layer positions and makes layer-to-layer matching unstable; the paper reports that this can diverge. A prediction-style intermediate loss was tested as a workaround for gradual stacking, but it did not solve the issue overall and could even degrade ASR and SF. The loss-weight study further reports that adding the intermediate loss improves performance substantially, with \(w=0.5\) chosen as the default because it gives the best average dev-set results [2606.11766].

## 5. Experimental validation on SUPERB

The empirical evaluation uses HuBERT base as teacher and a 12-layer Transformer student with 26.87M parameters. The student is trained on LibriSpeech 960h with AdamW, learning rate \(5\times10^{-4}\), batch size 64, weight decay \(10^{-4}\), for 75 epochs, and evaluated on four SUPERB tasks: phoneme recognition, ASR, slot filling, and speaker identification. Stagewise stacking uses \(B=4\) stages, so the model expands by 3 layers per stage. Two schedules are tested: equal scheduling and prop-1 scheduling, with speedup measured by wall-clock time [2606.11766].

The reported pattern is that stacking does accelerate training, but naive stacking harms downstream performance. Gradual stacking and MIDAS both lag behind non-stacking baselines on several tasks. Under equal scheduling, interleaved stacking reaches a 1.24× speedup and achieves 9.08 PER on PR, compared with 11.50 for gradual stacking and 10.75 for MIDAS; it also performs strongly on ASR, SF, and SID. Under prop-1 scheduling, it gets a 1.16× speedup and reaches 8.88 PER on PR, 9.99 WER on ASR, 28.45 CER on SF, and 73.60 accuracy on SID [2606.11766].

The paper emphasizes that these results are not only better than the stacking baselines, but in some cases competitive with or even better than the full-training model without stacking. That claim matters because it suggests that the benefit is not merely reduced computation in early stages. Instead, the schedule appears to preserve useful representation structure while also lowering training cost.

## 6. Layer similarity, representation hierarchy, and interpretation

The layer-similarity analysis is used to support the method’s intuition. The paper observes that adjacent layers in the distilled speech foundation model are more similar than distant ones. This makes adjacency-preserving duplication plausible: if copied layers are inserted next to their originals, they begin from positions where local representational similarity is already relatively high [2606.11766].

Interleaved stacking yields clearer block-like similarity structures and suggests that copied layers continue to serve similar roles after training. This does not imply that copied layers remain identical; rather, the preserved neighborhood appears to maintain a stable functional context during optimization. The paper’s broader interpretation is that layers in speech foundation models are not generic containers for depth, but carriers of depth-specific knowledge. From that perspective, interleaved stacking is an architectural schedule for conserving representation hierarchy during accelerated distillation, not simply a duplication heuristic [2606.11766].

An objective reading of the evidence therefore supports two distinctions. First, training acceleration and representational preservation are separable design goals. Second, intermediate-level KD is not automatically compatible with all stacking schemes; its stability depends on whether the student’s layer indexing remains semantically meaningful across stages.

## 7. Terminological scope and related uses of interleaving and stacking

The phrase “interleaved stacking” is domain-specific rather than universal. In the speech-distillation setting, it denotes in-place layer duplication during stagewise depth growth. Other fields use the terms “interleaving” and “stacking” differently, and these usages are technically distinct [2606.11766].

In interferometric H\,\textsc{i} analysis, “cubelet stacking” refers to stacking dirty image cubelets and PSF cubelets in the image domain before deconvolution. The essential point there is that image-domain stacking happens first and deconvolution happens only after the stack has increased the effective S/N; this is described as especially useful for extended or marginally resolved sources with interferometers [2101.06928].

In coding theory, high-order interleaved sum-rank-metric codes are formed by stacking \(s\) codewords of the same constituent code vertically. There, interleaving creates redundancy across rows so that a Metzner–Kapturowski-like decoder can recover errors of sum-rank weight \(t\) when \(t\le d-2\), \(s\ge t\), and \(\operatorname{rk}_{q^m}(E)=t\) [2303.17454].

In batched network coding, intrablock interleaving denotes packet-order optimization within a block under adaptive recoding. Its purpose is to separate packets of the same batch as much as possible while preserving bounded buffer size and bounded latency, unlike stream interleaving, which can make those quantities unbounded or hard to control [2105.07609].

These comparisons clarify a possible ambiguity. Across fields, “interleaving” generally refers to preserving or exploiting structure under some ordering constraint, and “stacking” refers to some form of aggregation or composition. In the specific speech-distillation method, however, the distinctive contribution is the preservation of layer position during depth expansion. That positional preservation, rather than interleaving in a generic sense, is the defining content of interleaved stacking [2606.11766].

Source: https://www.emergentmind.com/topics/interleaved-stacking