Papers
Topics
Authors
Recent
Search
2000 character limit reached

Incremental FastPitch: Real-Time TTS

Updated 12 December 2025
  • Incremental FastPitch is a text-to-speech model variant that uses chunk-based transformer decoders for incremental, low-latency speech synthesis.
  • It employs advanced attention masking and state-caching to maintain high speech fidelity while dramatically reducing end-to-end latency.
  • The architecture scales linearly with utterance length, making it ideal for real-time streaming and interactive TTS applications.

Incremental FastPitch is a variant of the FastPitch text-to-speech (TTS) architecture that enables low-latency, high-quality, incremental speech synthesis by re-architecting the decoder to operate on fixed-size Mel-spectrogram chunks using transformer-style chunk-based Feed-Forward Transformer (FFT) blocks, advanced attention masking for receptive-field control, and a state-caching inference regime. The method delivers real-time response suitable for streaming and interactive speech applications, cutting end-to-end latency dramatically compared to fully parallel transformer decoders while preserving output speech fidelity (Du et al., 2024).

1. Architectural Modifications in Incremental FastPitch

The canonical FastPitch system synthesizes speech in a fully parallel manner using transformer FFT blocks that process the upsampled encoder features uˉ∈RT×d\bar u \in \mathbb{R}^{T \times d} for the entire utterance in a single pass, yielding the full Mel-spectrogram Y∈RT×80Y \in \mathbb{R}^{T \times 80}. Incremental FastPitch preserves the encoder, duration, pitch, and energy predictors from FastPitch, but restructures the decoder to operate on non-overlapping chunks uˉ=[uˉ1;… ;uˉN]\bar u = [\bar u_1; \dots; \bar u_N], where each uˉi∈RSc×d\bar u_i \in \mathbb{R}^{S_c \times d} and ScS_c is the chunk size.

Each decoder layer comprises a stack of chunk-based FFT blocks, each with:

  • Multi-Head Attention (MHA) Sub-layer: At chunk tt, MHA computes query-key-value attention where keys and values are concatenated from a fixed-size cache of previous keys/values (pkit−1pk_i^{t-1}, pvit−1pv_i^{t-1}, size SpS_p) and the current chunk (uˉi\bar u_i projected by Y∈RT×80Y \in \mathbb{R}^{T \times 80}0, Y∈RT×80Y \in \mathbb{R}^{T \times 80}1). This design ensures the per-chunk inference cost is independent of chunk index Y∈RT×80Y \in \mathbb{R}^{T \times 80}2 or global sequence length Y∈RT×80Y \in \mathbb{R}^{T \times 80}3 and supports Mel-continuity.
  • Two-Layer Causal Convolutional Feed-Forward Network (FFN): Each FFN layer caches its last Y∈RT×80Y \in \mathbb{R}^{T \times 80}4 states (Y∈RT×80Y \in \mathbb{R}^{T \times 80}5 for Y∈RT×80Y \in \mathbb{R}^{T \times 80}6) to support causality over the chunk sequence.

The chunk-based FFT computational diagram can be summarized as:

pvit−1pv_i^{t-1}5

Mathematically, for MHA in head Y∈RT×80Y \in \mathbb{R}^{T \times 80}7, chunk Y∈RT×80Y \in \mathbb{R}^{T \times 80}8:

  • Y∈RT×80Y \in \mathbb{R}^{T \times 80}9
  • uˉ=[uˉ1;… ;uˉN]\bar u = [\bar u_1; \dots; \bar u_N]0
  • uˉ=[uˉ1;… ;uˉN]\bar u = [\bar u_1; \dots; \bar u_N]1
  • uˉ=[uˉ1;… ;uˉN]\bar u = [\bar u_1; \dots; \bar u_N]2
  • uˉ=[uˉ1;… ;uˉN]\bar u = [\bar u_1; \dots; \bar u_N]3; uˉ=[uˉ1;… ;uˉN]\bar u = [\bar u_1; \dots; \bar u_N]4

The causal-conv FFN processes similarly, caching and updating past conv states.

2. Training with Receptive-Field-Constrained Chunk Attention Masks

To ensure the model operates within its inference constraints during training, a strictly block-diagonal attention mask uˉ=[uˉ1;… ;uˉN]\bar u = [\bar u_1; \dots; \bar u_N]5 is applied in each decoder layer. The local receptive window per chunk is uˉ=[uˉ1;… ;uˉN]\bar u = [\bar u_1; \dots; \bar u_N]6; positions uˉ=[uˉ1;… ;uˉN]\bar u = [\bar u_1; \dots; \bar u_N]7 refer to the past cache and uˉ=[uˉ1;… ;uˉN]\bar u = [\bar u_1; \dots; \bar u_N]8 are the current chunk. The attention mask enforces:

uˉ=[uˉ1;… ;uˉN]\bar u = [\bar u_1; \dots; \bar u_N]9

Two regimes are explored:

  • Static Mask: uˉi∈RSc×d\bar u_i \in \mathbb{R}^{S_c \times d}0 fixed through training.
  • Dynamic Mask: uˉi∈RSc×d\bar u_i \in \mathbb{R}^{S_c \times d}1 randomly sampled per batch (e.g., uˉi∈RSc×d\bar u_i \in \mathbb{R}^{S_c \times d}2), enabling generalization to varying chunk configurations.

3. State-Caching Inference Algorithm

The inference algorithm maintains, per decoder layer uˉi∈RSc×d\bar u_i \in \mathbb{R}^{S_c \times d}3, fixed-size past-key/value and past-convolution caches for each processed chunk. For an input of uˉi∈RSc×d\bar u_i \in \mathbb{R}^{S_c \times d}4 chunks:

  1. For each chunk uˉi∈RSc×d\bar u_i \in \mathbb{R}^{S_c \times d}5, extract uˉi∈RSc×d\bar u_i \in \mathbb{R}^{S_c \times d}6.
  2. For every decoder layer:
    • Concatenate past caches with the current chunk for MHA.
    • Update caches using uˉi∈RSc×d\bar u_i \in \mathbb{R}^{S_c \times d}7 operations post-attention and post-convolution.
    • Normalize and propagate to the next block.
  3. Project final output to yield Mel chunk uˉi∈RSc×d\bar u_i \in \mathbb{R}^{S_c \times d}8 and emit for audio synthesis.

Per-chunk runtime is uˉi∈RSc×d\bar u_i \in \mathbb{R}^{S_c \times d}9, and the total cost scales linearly with the number of chunks (ScS_c0); inference cost per chunk is independent of chunk index.

4. Training Objectives and Optimization

The multi-loss objective is retained from FastPitch:

ScS_c1

with:

  • ScS_c2,
  • ScS_c3,
  • ScS_c4,
  • ScS_c5.

No extra regularization or loss terms are used; the chunked receptive field is enforced purely by the attention mask.

5. Experimental Results: Speech Quality and Latency

Empirical evaluation demonstrates Incremental FastPitch produces speech of comparable quality to parallel FastPitch at substantially reduced latency.

Table 1: Comparative Metrics

Model MOS (ScS_c6 95% CI) Latency (ms) RTF
Parallel FastPitch 4.185 ScS_c7 0.043 125.8 0.029
Inc. FastPitch (Static Mask) 4.178 ScS_c8 0.047 30.4 0.045
Inc. FastPitch (Dynamic Mask) 4.145 ScS_c9 0.052 30.4 0.045
Ground Truth 4.545 tt0 0.039 — —
  • Mel-Spectrogram Distance: Lowest when tt1 match at train/test; dynamic masking generalizes but with an tt28% higher MSD.
  • MOS: Incremental FastPitch matches FastPitch (MOS tt3 4.18) while cutting end-to-end latency by tt44x (from 125.8 ms to 30.4 ms).
  • RTF: Remains well above real-time requirements (tt51/22).

6. Computational Complexity and Latency Profile

Let tt6 = number of decoder layers, tt7 = chunk size, tt8 = cache size, tt9 = model dimension, pkit−1pk_i^{t-1}0 = attention heads, pkit−1pk_i^{t-1}1, pkit−1pk_i^{t-1}2 = FFN hidden dimension, pkit−1pk_i^{t-1}3 = conv kernel sizes.

  • MHA Cost/Chunk: pkit−1pk_i^{t-1}4
  • Conv-FFN Cost/Chunk: pkit−1pk_i^{t-1}5
  • Total Per-Chunk: pkit−1pk_i^{t-1}6

These costs are constant with respect to total utterance length (pkit−1pk_i^{t-1}7); only the number of chunks and chunk size matter. Start-up latency is determined by pkit−1pk_i^{t-1}8. With small pkit−1pk_i^{t-1}9 (e.g., pvit−1pv_i^{t-1}0 frames, pvit−1pv_i^{t-1}1 ms at pvit−1pv_i^{t-1}2 kHz and hop pvit−1pv_i^{t-1}3), sub-pvit−1pv_i^{t-1}4 ms startup latency is achievable.

7. Significance and Applications

Incremental FastPitch integrates chunked processing, cache-based attention, and tailored attention masking to deliver high MOS, low-latency, and scalable synthesis. Its design enables real-time, streaming, and interactive TTS workloads with reduced response times and memory requirements, while retaining the benefits of transformer-based parallel TTS models. The architecture provides a framework for incremental speech synthesis with complexity governed by chunk parameters rather than utterance length, pointing to new possibilities for low-latency, high-continuity TTS systems (Du et al., 2024).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Incremental FastPitch.