---
title: 'SPAR-K: Accelerated Inference for SLMs'
url: https://www.emergentmind.com/topics/spar-k
type: topic
---

# SPAR-K: Accelerated Inference for SLMs

SPAR-K is an inference-time acceleration framework for interleaved spoken language models (SLMs) in which text and speech tokens are generated within a single autoregressive stream. The method, introduced as “SPAR-K: Scheduled Periodic Alternating Early Exit for Spoken Language Models,” defines a modality-aware early-exit policy that applies shallow decoding primarily to speech tokens while preserving full-depth decoding for text and inserting periodic full-depth “refresh” steps to control autoregressive drift. In the reported experiments, SPAR-K largely preserves question-answering accuracy, with a maximum accuracy drop of \(0.82\%\), while reducing average speech decoding depth by up to \(11\%\) on Step-Audio-2-mini and \(5\%\) on GLM-4-Voice, with negligible changes in MOS and WER and no auxiliary computation overhead [2603.09215].

## 1. Conceptual basis and scope

SPAR-K expands to **Scheduled Periodic Alternating Early Exit**. Its target setting is the interleaved SLM, where text tokens and discrete speech tokens are emitted in a fixed alternation pattern. This model class is attractive because text tokens provide explicit semantic guidance for subsequent speech tokens and enable streaming speech synthesis, but it is also computationally expensive because speech is represented as long sequences of high-rate discrete units. In this setting, decoding every token at full transformer depth imposes a disproportionate cost on speech generation.

The method is explicitly **modality-aware**. Rather than treating all tokens with a single early-exit policy, it assumes that speech and text have different tolerances to shallow decoding. The paper’s argument is that intermediate-layer speech predictions can differ substantially from final-layer predictions while still synthesizing to perceptually similar audio, whereas intermediate-layer text predictions do not retain coherent sentence-level semantics. This suggests that the dominant acceleration opportunity in interleaved SLMs lies on the speech side rather than in a uniform token-agnostic exit rule [2603.09215].

## 2. Scheduled periodic alternating early exit

Let \(L\) be the total number of transformer layers and let \(h_t^{(\ell)} \in \mathbb{R}^d\) denote the hidden state at generation step \(t\) and layer \(\ell\). SPAR-K chooses an intermediate exit layer \(\ell_{EE}<L\) and applies it to most speech-token positions, while retaining full-depth decoding at periodic positions. The method is therefore schedule-based rather than confidence-routed.

Within each chunk of \(N_{\rm speech}\) speech tokens, SPAR-K partitions tokens into subgroups of size \(K\), decodes one position in each subgroup at full depth \(L\), and decodes the remaining \(K-1\) positions at \(\ell_{EE}\). The paper evaluates three concrete schedules:

1. **Even schedule** (\(K=2\)):
   \[
   \{L,\ell_{EE},L,\ell_{EE},\cdots\}
   \]

2. **Odd schedule** (\(K=2\)):
   \[
   \{\ell_{EE},L,\ell_{EE},L,\cdots\}
   \]

3. **Triple schedule** (\(K=3\)):
   \[
   \{L,\ell_{EE},\ell_{EE},L,\ell_{EE},\ell_{EE},\cdots\}
   \]

The motivation is control of **distribution shift** during autoregressive decoding. If a token is generated from an intermediate layer and then fed back as context, subsequent predictions are conditioned on a history that differs from the one the full model would have produced. A naive fixed-layer policy compounds that mismatch at every speech step. SPAR-K’s periodic full-depth positions act as refresh steps that restore the generation trajectory closer to the full-depth distribution often enough to prevent quality collapse while still reducing computation across the surrounding speech positions [2603.09215].

## 3. Layer-specific prediction heads and systems implementation

SPAR-K requires a layer-specific LM head \(g_\ell\) so that intermediate hidden states can directly predict next-token distributions:
\[
p_\ell(\cdot \mid y_{<t})=\operatorname{softmax}\left(g_\ell(h_t^{(\ell)})\right).
\]
The SLM backbone remains frozen; only the layer-specific head is trained. Training is performed by distilling the final-layer distribution into intermediate layers using cross-entropy over prefixes drawn from \(\mathcal D\). The paper states that the heads are trained on 18K instances from VoiceAssistant-400K, with pseudo-label distributions obtained from full-depth autoregressive generation.

A practical issue in early exit is that upper transformer layers do not produce KV-cache entries for positions that stop at \(\ell_{EE}\). SPAR-K addresses this with its periodic structure. When a later speech position is decoded at full depth, the model can simultaneously generate the missing higher-layer KV-cache for a preceding early-exit position, analogous to a miniature prefill over those local positions. The paper emphasizes that this avoids sacrificing decoding latency and is one reason the schedule-based design does not require auxiliary routing networks, confidence calibration, or per-step scoring overhead at inference time [2603.09215].

## 4. Evaluation protocol and main empirical results

The method is evaluated on two interleaved SLM backbones. **Step-Audio-2-Mini** has 28 layers and a text:speech interleaving ratio of \(1:4\). **GLM-4-Voice** has 40 layers and a text:speech interleaving ratio of \(13:26\). The evaluation uses four English datasets—AlpacaEval, Llama Questions, TriviaQA, and WebQuestion—described as spanning reasoning, factual QA, and dialogue. For the three QA datasets, answer accuracy is measured with GPT-4o-mini against ground truth; for AlpacaEval, an LLM-as-a-judge score normalized to 0–100 is used. Crucially, the reported evaluations are performed on ASR transcriptions of synthesized speech, not only on the text token stream.

Speech quality is measured with **MOS** from UTMOS-v2 and **ASR-WER**, where synthesized speech is transcribed by Whisper-large-v3 and compared with the model’s own interleaved text tokens. Efficiency is reported as average exit layer for text and speech positions, rather than wall-clock latency.

On **Step-Audio-2**, the baseline has mean score \(54.22\), MOS \(3.710\), ASR-WER \(1.51\), and text and speech exit layer 28. The highlighted SPAR-K setting is **Triple (22)**, which yields AlpacaEval \(46.33\), LlamaQA \(67.33\), TriviaQA \(51.00\), WebQA \(56.50\), mean \(55.29\), MOS \(3.668\), ASR-WER \(1.51\), text exit layer 28, and speech exit layer 25. This corresponds to an \(11\%\) speech layer reduction with no ASR-WER increase.

On **GLM-4-Voice**, the baseline has mean score \(52.37\), MOS \(2.982\), ASR-WER \(4.31\), and text and speech exit layer 40. The paper identifies **Even (36)** as the best overall balance: AlpacaEval \(50.70\), LlamaQA \(60.00\), TriviaQA \(42.00\), WebQA \(50.60\), mean \(50.83\), MOS \(2.950\), ASR-WER \(5.36\), text exit layer 40, and speech exit layer 38. This gives a \(5\%\) speech layer reduction with modest changes in speech quality. Across retained SPAR-K settings, the paper summarizes the largest average accuracy loss as \(0.82\%\) [2603.09215].

## 5. Empirical contrast with fixed-layer and confidence-based early exit

A central result is that **periodicity matters more than simply choosing a shallower layer**. On Step-Audio-2, the fixed-layer baseline **S2 Fixed (25)** has the same average speech exit layer as the SPAR-K schedules, but mean score drops to \(50.13\), MOS to \(3.058\), and ASR-WER rises to \(3.40\). On GLM-4-Voice, **G2 Fixed (38)** likewise shares the same \(5\%\) speech depth reduction as the retained SPAR-K settings, but yields mean \(51.48\), MOS \(2.662\), and ASR-WER \(43.60\). The paper interprets this as evidence that periodic refreshes are essential to prevent runaway autoregressive degradation such as redundant spoken tails and failure to terminate correctly.

The paper also argues that **confidence-based early exit**, common in text LLMs, is suboptimal for SLMs. On Step-Audio-2, the confidence baseline reduces average speech depth from 28 to 26.07 but mean task score drops from \(54.22\) to \(41.51\), MOS from \(3.710\) to \(1.651\), and ASR-WER rises from \(1.51\) to \(11.01\). On GLM-4-Voice, the confidence baseline is more viable but still inferior in practice: mean score \(52.89\) versus baseline \(52.37\), MOS \(2.866\) versus \(2.982\), ASR-WER \(7.62\) versus \(4.31\), with average speech exit depth \(37.03\) versus 40. The reported explanation is that speech tokens are acoustically local and redundant, so token-level confidence is a weaker proxy for perceptual fidelity than it is for text semantics [2603.09215].

The paper’s oracle-style analysis on GLM-4-Voice illustrates this asymmetry:

| Layer | MOS | Agreement with final layer |
|---|---:|---:|
| 16 | 2.922 | 19.14% |
| 21 | 2.905 | 22.13% |
| 26 | 2.978 | 32.83% |
| 31 | 2.985 | 39.57% |
| 36 | 2.996 | 46.21% |
| 40 | 3.004 | 100% |

MOS remains nearly constant even when agreement with the final layer is low. This suggests that large token-level divergence in speech space does not necessarily imply perceptually poor audio, weakening the premise of entropy-threshold exit rules imported from text-only models.

## 6. Limitations, boundary conditions, and research significance

SPAR-K’s gains are meaningful but not large in absolute depth reduction: the reported range is up to \(11\%\) on Step-Audio-2 and up to \(5\%\) on GLM-4-Voice rather than an order-of-magnitude acceleration. The method is also explicitly scoped to **interleaved** SLMs with fixed text/speech alternation and periodic speech chunks. The paper does not study non-interleaved spoken models, does not report end-to-end wall-clock latency or throughput, and does not provide a general theoretical criterion for choosing \(K\) and \(\ell_{EE}\); both are selected empirically on a held-out set using WER and MOS.

A further limitation is that **text-token acceleration remains unsolved** within the framework. On GLM-4-Voice, applying SPAR-K to text tokens alone causes severe semantic collapse: AlpacaEval \(10.00\), LlamaQA \(42.00\), TriviaQA \(0.00\), WebQA \(21.70\), mean \(18.43\), MOS \(2.979\), and ASR-WER \(12.80\). A hybrid system combining confidence-based text exit with SPAR-K speech exit still produces a notable mean drop to \(48.61\). This indicates that the modality-aware premise is not merely an implementation detail but the central constraint of the method.

The broader significance of SPAR-K is therefore methodological rather than architectural. It shows that acceleration for spoken language models should be designed around the asymmetric statistical roles of text and speech tokens. A plausible implication is that future SLM inference systems will need heterogeneous policies across modalities, with text treated as semantically fragile and speech treated as locally redundant but susceptible to on-policy drift unless periodically corrected. In that sense, SPAR-K defines a specialized early-exit regime for interleaved SLMs rather than a general-purpose transformer acceleration rule [2603.09215].

Source: https://www.emergentmind.com/topics/spar-k