---
title: Masked Chunking in ASR
url: https://www.emergentmind.com/topics/masked-chunking-in-asr
type: topic
---

# Masked Chunking in ASR

Masked chunking in automatic speech recognition (ASR) refers to a family of architectural, inference, and batching strategies that segment feature or hypothesis sequences into distinct "chunks," apply masking to structure computation or loss, and restrict attention or decoding to subsets of the data. Masked chunking enables efficient streaming, non-autoregressive, and long-form processing in state-of-the-art ASR systems, minimizing both latency and resource usage while maintaining recognition accuracy. Implementations span CTC-based mask-refinement loops, chunk-aware self-attention, and explicit batch-level masking for scaling to arbitrarily long or highly variable audio inputs.

## 1. Formal Definitions and Theoretical Underpinnings

Masked chunking operates by partitioning the input feature sequence, intermediate representations, or output hypotheses into discrete segments—referred to as "chunks"—and then applying masks to restrict self-attention, convolution, or decoding steps to within or between these chunks. In contemporary ASR, this manifests in several primary forms:

- **Chunked Attention Masks:** Let $X\in\mathbb{R}^{T\times d}$ be $T$ frames of input features. For chunk size $C_\ell$ at layer $\ell$, chunked masking builds a binary matrix $A^{(\ell)}_{\mathrm{chunked}}$ with
  $$
  A^{(\ell)}_{\mathrm{chunked}}(i,j) =
  \begin{cases}
    1, & \left\lfloor \frac{i}{C_\ell} \right\rfloor = \left\lfloor \frac{j}{C_\ell} \right\rfloor \\
    0, & \text{otherwise}
  \end{cases}
  $$
  All queries attend freely within their chunk and are strictly causal across chunk boundaries [2211.01438].

- **Mask-and-Refine Loop:** In Mask-CTC, the output hypothesis $\hat{Y}$ is decomposed into observed (high-confidence) and masked (low-confidence) token "chunks," where the mask $M$ is defined by token-level CTC confidence falling below $P_{\mathrm{thres}}$ [2005.08700].

- **Masked Batching:** For efficient batched processing, especially with disparate audio lengths, binary chunk-position masks $M\in\{0,1\}^{M\times(c + l + r)}$ identify valid frames post-chunking and relative context augmentation, and are applied to all convolutional and attention modules [2502.14673].

This theoretical scaffolding enables sublinear scaling of memory and time, critical for streaming, long-form, and low-latency ASR deployment.

## 2. Model Architectures Incorporating Masked Chunking

Masked chunking has catalyzed several architectural variants tailored for different constraints:

- **Mask-CTC:** A dual-headed model with a CTC-based Transformer encoder and a non-autoregressive CMLM decoder. Here, masked chunking occurs in hypothesis space, not time or feature space. The mask-and-refine loop iteratively distinguishes "observed" from "masked" tokens, and only re-predicts the masked chunk(s) via the decoder [2005.08700].

- **Chunk-Aware Self-Attention:** In SCAMA and LC-SAN-M, the encoder decomposes inputs into chunks of $c$ frames, applies self-attention restricted to current and past chunks (LC-SAN-M), and maintains separate chunk boundaries for streaming control. The decoder imposes masks at chunk boundaries and leverages a jointly trained predictor to determine output emission count per chunk [2006.01712].

- **ChunkFormer Backbone:** The encoder is split into chunkwise, overlapping windows with explicit left and relative right context at every layer. Masked chunking is applied across the entire forward pass: convolutions zero out invalid frames, and attention logits are $-1e9$ masked wherever the binary position mask indicates padding or overlap [2502.14673].

- **Transformer-Transducer with Variable Masking:** Variable attention masking enables a single model to generalize across fixed, chunked, and variable masking regimes. Chunked masking allows full self-attention within each chunk but zeroes attention between chunks, formalized for each layer and switched via sampled mask configurations at training [2211.01438].

## 3. Masked Chunking Algorithms and Inference Procedures

Algorithmic implementation of masked chunking differs by paradigm but typically comprises:

- **Chunk Boundary and Mask Construction:** For chunk-based self-attention, chunk boundaries are defined by $k(i) = \lfloor i / C_\ell\rfloor$. Binary masks are constructed per layer; for each $(i, j)$, $M[i, j] = 1$ iff queries and keys are in same chunk.

- **Mask-and-Refine Decoding:** In Mask-CTC, greedy CTC decoding yields initial sequence $\hat{Y}$. Token set $M = \{l\mid \hat{P}(\hat{y}_l\mid X) < P_\mathrm{thres}\}$ is masked. The CMLM decoder predicts masked positions, either in one pass or K "easy-first" refinement iterations, with updated confidence at each step [2005.08700].

- **Chunk-Aware Streaming Decoding:** In SCAMA, for each encoder chunk, the predictor outputs number of tokens $\hat{y}_k$; decoder attends exclusively to the first $k\cdot c$ encoder frames, enforced via per-query mask $M_{\ell, j} = 0$ if $j\leq k\cdot c$ and $-\infty$ otherwise [2006.01712].

- **Masked Batching for Variable-Length Inputs:** Prior to attention or convolution, the input batch is reshaped into uniform chunked windows, augmented with enough left/right context. Binary mask $M$ identifies valid positions across all utterances and is reused layerwise, eliminating any padding inefficiency and preventing spurious context bleed [2502.14673].

Pseudocode in the referenced works formalizes these procedures; e.g., Mask-CTC pseudocode enumerates the mask-and-refine iteration over hypothesis tokens [2005.08700].

## 4. Empirical Findings and Benchmarks

Substantial empirical validation has shown masked chunking unlocks favorable tradeoffs in ASR:

| System / Setting        | WER (%)          | Latency [RTF/ms]       | Dataset                  |
|------------------------|------------------|------------------------|--------------------------|
| CTC only (1 iter) [2005.08700] | 17.9 (WSJ)        | 0.03 (RTF)             | WSJ eval92               |
| Mask-CTC (10 iter)     | 12.1             | 0.07                   | WSJ eval92               |
| CTC-attn AR (greedy)   | 11.3             | 0.97                   | WSJ eval92               |
| LC-SAN-M+SCAMA         | 7.39 (CER, 600ms)| 600 ms (encoder chunk)  | AISHELL-1                |
| ChunkFormer (masked)   | 16.60–18.36      | 0.8 s (batch time)      | LibriSpeech, Earnings-21 |
| TT with chunked mask   | 3.62 (WER)       | 453 ms (PRWL)           | US English 60h test      |

- **Latency vs. accuracy:** Chunked masking reduces partial-result word latency (PRWL) by nearly $2\times$ relative to fixed masking at only $\sim0.1\%$ absolute WER cost [2211.01438].
- **Long-form scaling:** ChunkFormer's masked chunking enables transcription of up to 16 hours on a 80GB GPU, with $\sim7.7\%$ absolute WER reductions on long-form tasks relative to previous baselines, and $3\times$–$4\times$ RAM/time savings in multi-length batching [2502.14673].
- **Iterative refinement:** Mask-CTC with $5$–$10$ iterations closes over $80\%$ of the WER gap between vanilla CTC and AR models, with $10\times$–$20\times$ faster inference [2005.08700].
- **Variable masking for rescoring:** Variable masking allows transformer-transducer models to be deployed seamlessly in both streaming (small chunk) and second-pass (large chunk) rescoring, yielding up to $8\%$ relative WER reduction [2211.01438].
- **Predictor stability:** In SCAMA, joint predictor training yields more stable chunk transitions and is robust to large-channel, industrial-scale Mandarin data [2006.01712].

## 5. Practical Implementation Guidelines

Optimal configuration of masked chunking is scenario-dependent:

- **Chunk size:** Small (60–120 ms) for lowest streaming latency, at minor WER cost; medium (180–240 ms) as the accuracy/latency "sweet spot." Very large chunks restore full-context but forfeit latency advantage [2211.01438].
- **Context windows:** Minimal history (0.72–2.0 s left context) suffices for streaming; full-past for rescoring [2211.01438]. Relative right context per layer is cumulative in deep chunkwise architectures [2502.14673].
- **Mask sampling:** Uniform sampling over a discrete mask set (chunks/left context) achieves configurable deployment with limited mode collapse [2211.01438].
- **Batching:** Masked batching eliminates padding overhead and maximizes hardware utilization; position masks are precomputed and broadcast to all convolutional/attention heads [2502.14673].
- **Losses:** Joint CTC and AED/RNN-T losses are standard. For models with chunkwise predictors, predictor cross-entropy loss is critical [2006.01712].

A plausible implication is that maintaining mask flexibility at runtime strongly favors reusable, hardware-efficient, and latency-aware ASR infrastructure.

## 6. Comparative Analysis and Limitations

Masked chunking strategies outperform prior methods under diverse accuracy, scalability, and deployment constraints:

- **Versus fixed look-ahead masking:** Chunked masking matches accuracy while often halving PRWL; variable attention masking further bridges the gap between streaming and offline modes [2211.01438].
- **Versus MoChA/monotonic attention:** SCAMA with a learned predictor is more stable and more parallelizable, with lower absolute CER loss under tight latency [2006.01712].
- **Versus naive batching:** Masked batch chunking avoids $O(L^2)$ attention costs, eliminates "fake" padding, and enables seamless mixing of long/short utterances in live serving [2502.14673].
- **Limitations:** Chunk size, relative right context, and mask density must be tuned to avoid efficiency loss or context fragmentation. Masked chunking in concatenated or highly discontinuous input still necessitates careful masking logic to prevent information leakage between utterances.

This suggests that further advances will focus on mask learning, adaptive chunking, and dynamically configurable architectures to handle diverse ASR scenarios.

Source: https://www.emergentmind.com/topics/masked-chunking-in-asr