---
title: SAQ-Decoder in SAC Speech Codec
url: https://www.emergentmind.com/topics/saq-decoder
type: topic
---

# SAQ-Decoder in SAC Speech Codec

Searching arXiv for recent papers mentioning “SAQ-Decoder” and the SAC codec paper to ground the article in the relevant literature.
SAQ-Decoder denotes the decoder component of SAC, a neural speech codec with semantic-acoustic dual-stream quantization, in which semantic and acoustic discrete token streams are fused and mapped back to a 16 kHz waveform. In the SAC formulation, the decoder is responsible not only for waveform reconstruction but also for auxiliary recovery of fine-grained continuous semantic features and estimation of speaker embeddings from the fused latent sequence. The design is motivated by the observation that existing speech codecs often struggle to balance high-quality reconstruction with semantically rich representations; SAC addresses this by disentangling semantic and acoustic modeling into two dedicated streams [2510.16841].

## 1. Definition and placement within SAC

Within SAC, the decoder operates downstream of two quantized streams. The semantic stream consists of quantized semantic embeddings $S_q'$ at 50 Hz. The original semantic features $S_c$ at 50 Hz are extracted by a frozen pre-trained speech tokenizer, then pooled to 12.5 Hz, quantized to tokens, embedded to $S_q$, and up-sampled back to 50 Hz via a small ConvNeXt adapter to produce $S_q'$. The acoustic stream consists of quantized acoustic embeddings $A_q$ at either 25 Hz in the low-bitrate setting or 50 Hz in the high-bitrate setting, obtained by an Encodec-style encoder with temporal downsampling strides $\tau=(2,2,4,5,8)$ or $\tau=(2,4,5,8)$ and single-codebook quantization [2510.16841].

The decoder therefore sits at the convergence point of SAC’s dual-stream representation. Its immediate input is not raw discrete indices, but continuous vectors recovered from those indices: $S_q'$ for semantics and $A_q$ for acoustics. This arrangement is central to the codec’s claim that semantic and acoustic factors can be optimized for distinct roles while remaining jointly decodable.

## 2. Decoder architecture

The decoder begins with a fusion prenet. The semantic and acoustic streams are concatenated along the channel dimension to form $U$, which is then passed through a ConvNeXt-based prenet. This prenet temporally upsamples to 50 Hz if needed and produces the fused feature sequence $F$ [2510.16841].

The main waveform decoder is a mirrored decoder relative to the Encodec-style encoder. The fused sequence $F$ is processed by a stack of convolutional layers interleaved with transposed-convolution upsampling layers. The deconvolution strides are $\tau=(8,5,4,2)$, exactly inverting the encoder’s downsampling so that the final output is a 16 kHz waveform $\tilde{x}$ [2510.16841].

Two auxiliary decoder-side heads are attached to the same fused representation. First, an auxiliary semantic decoder applies a small CNN to $F$ in order to reconstruct fine-grained continuous semantic features $\tilde{S}_c$ at 50 Hz. Second, an auxiliary speaker predictor summarizes $F$ using temporal mean and standard deviation,
$$
f=[\mathrm{Mean}_t(F);\mathrm{Std}_t(F)],
$$
and passes this statistic through a two-layer MLP, denoted “Proj,” to produce $\tilde{S}_p$, an estimate of the speaker embedding [2510.16841].

This architecture makes the fused representation multi-purpose: it drives waveform synthesis, semantic feature reconstruction, and speaker inference. A plausible implication is that SAC does not treat speech reconstruction as a purely signal-level inversion problem, but as a jointly constrained reconstruction problem in waveform, semantic, and speaker spaces.

## 3. Mathematical objectives and training signals

The waveform reconstruction loss is defined as a multi-scale $L_1$ objective on both linear- and log-scale spectrograms of $\tilde{x}$ versus $x$:
$$
L_{\mathrm{recon}}=\sum_{\mathrm{scales}}\|S_{\mathrm{lin}}(\tilde{x})-S_{\mathrm{lin}}(x)\|_1+\|S_{\mathrm{log}}(\tilde{x})-S_{\mathrm{log}}(x)\|_1.
$$
The paper notes that no closed-form equation is given in the text and that the formulation follows DAC [2510.16841].

For the acoustic stream, vector quantization is trained with
$$
L_{\mathrm{vq}}=\|\mathrm{sg}[a]-e\|_2^2+\beta\|a-\mathrm{sg}[e]\|_2^2,
$$
where $\beta$ is the commitment weight, fixed at $0.25$, $\mathrm{sg}[\cdot]$ is the stop-gradient operator, and an additional codebook loss weight $4.0$ scales the first term [2510.16841].

Adversarial supervision is provided by two discriminators, Multi-Period and Multi-Scale STFT, trained with least-squares GAN. The generator therefore also receives adversarial loss $L_{\mathrm{adv}}$ and feature-matching loss $L_{\mathrm{feat}}$, where the latter is an $L_1$ distance between discriminator feature maps [2510.16841].

The semantic auxiliary loss is
$$
L_{\mathrm{sem}}=\|\tilde{S}_c-S_c\|_2^2,
$$
with $\tilde{S}_c$ predicted by the CNN decoder and $S_c$ given by the frozen tokenizer’s continuous semantic features. The speaker auxiliary loss is
$$
\tilde{S}_p=\mathrm{Proj}(f), \qquad
L_{\mathrm{spk}}=\|\tilde{S}_p-S_p\|_2^2,
$$
where $S_p$ is the frozen ERes2Net speaker embedding [2510.16841].

These terms are combined into the generator objective
$$
\mathcal{L}_G=
\lambda_{\mathrm{recon}}L_{\mathrm{recon}}
+\lambda_{\mathrm{vq}}L_{\mathrm{vq}}
+\lambda_{\mathrm{adv}}L_{\mathrm{adv}}
+\lambda_{\mathrm{feat}}L_{\mathrm{feat}}
+\lambda_{\mathrm{sem}}L_{\mathrm{sem}}
+\lambda_{\mathrm{spk}}L_{\mathrm{spk}}.
$$
The reported weights are $\lambda_{\mathrm{recon}}=15$, $\lambda_{\mathrm{vq}}=1$, $\lambda_{\mathrm{adv}}=1$, $\lambda_{\mathrm{feat}}=2$, $\lambda_{\mathrm{sem}}=1000$, and $\lambda_{\mathrm{spk}}=10$ [2510.16841].

## 4. De-quantization, continuous reconstruction, and implementation parameters

The decoder’s de-quantization procedure is direct. In the semantic stream, each discrete token index $i$ selects the $i$-th row of the frozen semantic tokenizer’s codebook to produce embedding $S_q$, after which a ConvNeXt upsampler yields $S_q'$. In the acoustic stream, each frame’s continuous embedding $A$ is replaced by its nearest codebook entry $A_q$ through $L_2$ search in the learned single codebook of size $16\,384$. At decoding time, no further inverse quantization is needed, because $S_q'$ and $A_q$ are already continuous vectors; these are concatenated and passed into the prenet [2510.16841].

Several decoder-side hyperparameters characterize the implementation. The semantic codebook size is $16\,384$, and the acoustic codebook size is also $16\,384$. Semantic features are pooled from 50 Hz to 12.5 Hz before quantization and restored to 50 Hz after upsampling. Acoustic features operate at 25 Hz for low bitrates and 50 Hz for high bitrates. The generator has approximately $277$ million parameters, of which approximately $249$ million are trainable, while the semantic tokenizer and speaker encoder remain frozen [2510.16841].

Training uses AdamW with $\beta_1=0.8$, $\beta_2=0.9$, initial learning rate $10^{-4}$, and exponential decay. The generator is pretrained for $1\,500$ steps before the discriminator is activated, and EMA is used [2510.16841].

| Component | Reported specification |
|---|---|
| Prenet | ConvNeXt-based; exact number of layers not specified |
| Main decoder strides | $\tau=(8,5,4,2)$ |
| Semantic codebook | $16\,384$ |
| Acoustic codebook | $16\,384$ |
| Acoustic frame rate | 25 Hz or 50 Hz |
| Output | 16 kHz waveform $\tilde{x}$ |

A plausible implication is that the decoder is intentionally structured so that semantic and acoustic streams meet only after each has already been converted into a continuous latent trajectory, thereby simplifying the fusion problem to one of feature integration rather than symbolic sequence transduction.

## 5. Reconstruction quality and bitrate regimes

SAC evaluates reconstruction with STOI, PESQ-NB, PESQ-WB, UTMOS, Speaker SIM, and WER. At the high-bitrate setting, corresponding to a 62.5 Hz token rate and 875 bps, the reported scores are UTMOS $=4.25$, WER $=2.35\%$, SIM $=0.86$, PESQ-NB $=3.15$, PESQ-WB $=2.59$, and STOI $=0.93$. The paper states that this UTMOS is the best among approximately 1 kbps codecs and even above ground-truth $4.09$, while the WER is close to ground truth $2.16\%$ [2510.16841].

At the low-bitrate setting, corresponding to a 37.5 Hz token rate and 525 bps, the reported scores are UTMOS $=4.27$, WER $=2.53\%$, SIM $=0.78$, PESQ-NB $=2.74$, PESQ-WB $=2.18$, and STOI $=0.90$ [2510.16841].

Under noisy conditions on LibriSpeech test-other, SAC is reported to still top UTMOS, with values $3.84/3.90$, maintain low WER at approximately $5.8/6.4\%$, and retain SIM of approximately $0.77$–$0.85$ [2510.16841].

| Setting | Token rate / bitrate | Key reported outcomes |
|---|---|---|
| High bitrate | 62.5 Hz, 875 bps | UTMOS 4.25, WER 2.35%, SIM 0.86 |
| Low bitrate | 37.5 Hz, 525 bps | UTMOS 4.27, WER 2.53%, SIM 0.78 |
| Noisy test-other | not separately restated | UTMOS 3.84/3.90, WER $\sim$5.8/6.4%, SIM $\sim$0.77–0.85 |

These results are presented as evidence that the decoder preserves perceptual quality and intelligibility across bitrate regimes and in noisy conditions. Because the decoder is the component that fuses the semantic and acoustic streams back into waveform space, the reported quality metrics also function as an empirical validation of the decoder design itself.

## 6. Ablations, disentanglement, controllability, and nomenclature

The paper’s ablation studies directly probe the decoder’s role in disentanglement. Removing the speaker loss by setting $\lambda_{\mathrm{spk}}=0$ causes SIM to drop from $0.78$ to $0.65$, while PESQ increases slightly. Removing the semantic loss by setting $\lambda_{\mathrm{sem}}=0$ has negligible effect on STOI, PESQ, UTMOS, and WER, which the paper interprets as confirming that the frozen semantic stream plus the dual-stream design already preserves semantics [2510.16841].

Semantic-only and acoustic-only reconstructions further characterize the decoder. When the acoustic stream $A_q$ is masked, the semantic-only output yields WER $=3.99\%$ versus SemantiCodec’s $30.67\%$, SIM $=0.17$, and MSIM $=0.64$, which the paper describes as indicating strong lexical preservation and clean timbre removal. When the semantic stream $S_q'$ is masked, the acoustic-only output is noise-like and has zero intelligibility. Full reconstruction retains both high-fidelity timbre and semantics. The paper concludes that this confirms the decoder’s ability to fuse or ignore each stream cleanly [2510.16841].

These ablations suggest that the decoder is not merely combining two redundant latent codes. Rather, the semantic stream appears to carry lexical content in a form that remains intelligible even when speaker identity is largely removed, whereas the acoustic stream alone is insufficient for intelligible reconstruction. This suggests a decoder-level factorization of speech content and timbre that may be useful for controllable speech applications, a possibility explicitly noted in the SAC analysis [2510.16841].

The term “SAQ-Decoder” requires care because it is not unique across the 2025 arXiv literature. In the SAC paper, it refers to the decoder just described. However, the same label also appears in unrelated domains: as a simulated-annealing decoder for the XZZX code [2509.17837], and as a stabilizer-aware quantum error-correction decoder built from a dual-stream transformer and constraint-aware post-processing [2512.08914]. In speech codec literature, the relevant referent is the SAC decoder of semantic-acoustic dual-stream quantization [2510.16841].

Source: https://www.emergentmind.com/topics/saq-decoder