Papers
Topics
Authors
Recent
Search
2000 character limit reached

Interleaved Stacking in Speech Model Distillation

Updated 5 July 2026
  • Interleaved stacking is a stagewise method that duplicates layers immediately after their originals to preserve depth-specific roles in speech foundation models.
  • It addresses the limitation of layer relocation seen in traditional stacking methods, ensuring early, middle, and late layers retain distinct, semantic identities.
  • Experimental results on SUPERB tasks demonstrate that interleaved stacking accelerates training while maintaining or improving performance in low-resource distillation setups.

Interleaved stacking is a stagewise stacking method for speech foundation model distillation in which copied layers are inserted immediately after their source layers so that layer position is preserved as model depth increases. It was introduced to address a limitation of existing stacking methods: while shallow-to-deep growth reduces training cost, copying blocks to the top of the network repeatedly relocates layers and can degrade downstream performance in speech foundation models, where early, middle, and late layers encode distinct layer-specific knowledge. In this setting, interleaved stacking is not merely a schedule for reducing wall-clock time; it is a positional constraint on depth growth designed to preserve the semantic identity of layers during distillation (Kim et al., 10 Jun 2026).

1. Problem setting in speech foundation model distillation

The method was proposed in the context of distilling a large speech foundation model into an efficient student model for low-resource environments. Distillation reduces inference latency, but it requires additional student-model training, and the training efficiency of speech foundation model distillation had remained underexplored. The specific problem is therefore not whether distillation is useful, but how to accelerate student training without sacrificing the representational structure that makes distilled speech models effective (Kim et al., 10 Jun 2026).

The paper frames stacking as a natural candidate for training acceleration. If the target student has depth NN and training is split into BB stages, then a shallower model can be trained first and expanded progressively until the target depth is reached. Early stages are cheaper because they use fewer layers. However, in speech foundation models, this economy creates a structural risk: different layers specialize differently, and layer-specific knowledge matters. A stacking rule that changes the meaning of a layer’s depth position forces the model to relearn roles at each stage. This suggests that, in speech, speed-oriented depth growth must be constrained by representational continuity rather than treated as a purely architectural convenience (Kim et al., 10 Jun 2026).

2. Stagewise stacking and the problem of layer relocation

In standard stagewise stacking, K=N/BK=N/B, and at stage bb the model has bKbK layers. The network grows from KK to $2K$ to $3K$, and so on until NN. The efficiency benefit comes from training the shallower stages first. The difficulty lies in how new layers are initialized and placed (Kim et al., 10 Jun 2026).

Two prior strategies are described. Gradual stacking initializes the next stage by copying the last KK layers and stacking them on top: BB0 MIDAS instead copies an intermediate block: BB1 where BB2 (Kim et al., 10 Jun 2026).

The common limitation is that copied layers are relocated into a new structural context. A layer that was late in one stage may become middle in the next stage. In a speech foundation model, that relocation is problematic because early and late layers do not behave as interchangeable “depth slots.” The paper therefore treats positional instability, rather than copying itself, as the main source of performance degradation (Kim et al., 10 Jun 2026).

Method Copy strategy Positional effect
Gradual stacking Copy the last BB3 layers and stack them on top Late layers can become middle layers
MIDAS Copy an intermediate block Copied block is relocated into a new context
Interleaved stacking Copy evenly spaced layers and insert each after its source layer Relative layer position is preserved

A common misconception is that all stacking methods are equivalent up to training speed. The reported comparison rejects that view: the decisive variable is not only reduced computation in early stages, but whether depth growth preserves the representational role associated with layer position.

3. Interleaved stacking as positional preservation

Interleaved stacking differs from both gradual stacking and MIDAS by copying evenly spaced layers and inserting them immediately after their originals. In the BB4-th stage, it copies layer BB5 for BB6, so that each copied layer is duplicated in place: BB7 The paper emphasizes that the important point is not the algebraic identity by itself, but the placement rule: the copy stays adjacent to its source instead of being appended at the top (Kim et al., 10 Jun 2026).

Conceptually, this means that early, middle, and late layers continue to occupy roughly the same relative positions as the model expands. The copied layers are not “moved upward” to become a different part of the hierarchy. The method therefore preserves what the paper describes as layer identity across the entire depth-growth schedule. In speech foundation models, where layers encode different kinds of information at different depths, that consistency is presented as especially critical (Kim et al., 10 Jun 2026).

The central claim is thus narrower and more technical than a generic “better initialization” argument. Interleaved stacking is designed to preserve the mapping between depth position and function during stagewise expansion. A plausible implication is that the method is most relevant in settings where representational specialization across depth is strong, and less critical in architectures where layers are more nearly interchangeable.

4. Distillation objective and intermediate supervision

The method is embedded in a knowledge-distillation setup that combines output-level and intermediate-level supervision. The output-level KD loss is defined between the teacher output BB8 and the student output BB9, with a projection layer to match dimensions: K=N/BK=N/B0 The total objective adds an intermediate-level KD term: K=N/BK=N/B1 with K=N/BK=N/B2 controlling the strength of intermediate supervision (Kim et al., 10 Jun 2026).

For a K=N/BK=N/B3-layer initial-stage student, there are K=N/BK=N/B4 intermediate losses, and the teacher targets are fixed to evenly spaced intermediate layers. The paper gives a 12-layer teacher/student example with K=N/BK=N/B5, so K=N/BK=N/B6: teacher layers 4 and 8 may be compared to student layers 1 and 2 in stage 1, then to student layers 2 and 4 in stage 2. Importantly, the teacher target indices are kept fixed across stages, which the paper presents as stabilizing and speeding up training (Kim et al., 10 Jun 2026).

This interaction between stacking and distillation is a central part of the proposal. Interleaved stacking works well with intermediate-level KD losses because student layers remain in consistent positions across stages. Gradual stacking, by contrast, changes layer positions and makes layer-to-layer matching unstable; the paper reports that this can diverge. A prediction-style intermediate loss was tested as a workaround for gradual stacking, but it did not solve the issue overall and could even degrade ASR and SF. The loss-weight study further reports that adding the intermediate loss improves performance substantially, with K=N/BK=N/B7 chosen as the default because it gives the best average dev-set results (Kim et al., 10 Jun 2026).

5. Experimental validation on SUPERB

The empirical evaluation uses HuBERT base as teacher and a 12-layer Transformer student with 26.87M parameters. The student is trained on LibriSpeech 960h with AdamW, learning rate K=N/BK=N/B8, batch size 64, weight decay K=N/BK=N/B9, for 75 epochs, and evaluated on four SUPERB tasks: phoneme recognition, ASR, slot filling, and speaker identification. Stagewise stacking uses bb0 stages, so the model expands by 3 layers per stage. Two schedules are tested: equal scheduling and prop-1 scheduling, with speedup measured by wall-clock time (Kim et al., 10 Jun 2026).

The reported pattern is that stacking does accelerate training, but naive stacking harms downstream performance. Gradual stacking and MIDAS both lag behind non-stacking baselines on several tasks. Under equal scheduling, interleaved stacking reaches a 1.24× speedup and achieves 9.08 PER on PR, compared with 11.50 for gradual stacking and 10.75 for MIDAS; it also performs strongly on ASR, SF, and SID. Under prop-1 scheduling, it gets a 1.16× speedup and reaches 8.88 PER on PR, 9.99 WER on ASR, 28.45 CER on SF, and 73.60 accuracy on SID (Kim et al., 10 Jun 2026).

The paper emphasizes that these results are not only better than the stacking baselines, but in some cases competitive with or even better than the full-training model without stacking. That claim matters because it suggests that the benefit is not merely reduced computation in early stages. Instead, the schedule appears to preserve useful representation structure while also lowering training cost.

6. Layer similarity, representation hierarchy, and interpretation

The layer-similarity analysis is used to support the method’s intuition. The paper observes that adjacent layers in the distilled speech foundation model are more similar than distant ones. This makes adjacency-preserving duplication plausible: if copied layers are inserted next to their originals, they begin from positions where local representational similarity is already relatively high (Kim et al., 10 Jun 2026).

Interleaved stacking yields clearer block-like similarity structures and suggests that copied layers continue to serve similar roles after training. This does not imply that copied layers remain identical; rather, the preserved neighborhood appears to maintain a stable functional context during optimization. The paper’s broader interpretation is that layers in speech foundation models are not generic containers for depth, but carriers of depth-specific knowledge. From that perspective, interleaved stacking is an architectural schedule for conserving representation hierarchy during accelerated distillation, not simply a duplication heuristic (Kim et al., 10 Jun 2026).

An objective reading of the evidence therefore supports two distinctions. First, training acceleration and representational preservation are separable design goals. Second, intermediate-level KD is not automatically compatible with all stacking schemes; its stability depends on whether the student’s layer indexing remains semantically meaningful across stages.

The phrase “interleaved stacking” is domain-specific rather than universal. In the speech-distillation setting, it denotes in-place layer duplication during stagewise depth growth. Other fields use the terms “interleaving” and “stacking” differently, and these usages are technically distinct (Kim et al., 10 Jun 2026).

In interferometric H\,\textsc{i} analysis, “cubelet stacking” refers to stacking dirty image cubelets and PSF cubelets in the image domain before deconvolution. The essential point there is that image-domain stacking happens first and deconvolution happens only after the stack has increased the effective S/N; this is described as especially useful for extended or marginally resolved sources with interferometers (Chen et al., 2021).

In coding theory, high-order interleaved sum-rank-metric codes are formed by stacking bb1 codewords of the same constituent code vertically. There, interleaving creates redundancy across rows so that a Metzner–Kapturowski-like decoder can recover errors of sum-rank weight bb2 when bb3, bb4, and bb5 (Jerkovits et al., 2023).

In batched network coding, intrablock interleaving denotes packet-order optimization within a block under adaptive recoding. Its purpose is to separate packets of the same batch as much as possible while preserving bounded buffer size and bounded latency, unlike stream interleaving, which can make those quantities unbounded or hard to control (Yin et al., 2021).

These comparisons clarify a possible ambiguity. Across fields, “interleaving” generally refers to preserving or exploiting structure under some ordering constraint, and “stacking” refers to some form of aggregation or composition. In the specific speech-distillation method, however, the distinctive contribution is the preservation of layer position during depth expansion. That positional preservation, rather than interleaving in a generic sense, is the defining content of interleaved stacking (Kim et al., 10 Jun 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Interleaved Stacking.