Papers
Topics
Authors
Recent
Search
2000 character limit reached

Make-An-Audio 2: Advanced Text-to-Audio Generation

Updated 26 November 2025
  • Make-An-Audio 2 is a latent diffusion-based text-to-audio generation model that uses temporal parsing and dual text encoders to improve semantic alignment and variable-length audio synthesis.
  • It replaces conventional 2D U-Nets with a temporal transformer-based denoiser, enhancing temporal consistency and reducing artifacts in generated audio.
  • The model leverages LLM-guided data augmentation to expand training samples, achieving superior objective and subjective performance on standard T2A benchmarks.

Make-An-Audio 2 is a latent diffusion-based text-to-audio (T2A) generation model designed to address shortcomings in semantic alignment and temporal consistency that affect prior T2A systems. Traditional approaches, including those relying on 2D spatial structures (e.g., 2D U-Nets), often suffer from misaligned semantics and poor handling of variable-length audio, with limited modeling of temporal information. Make-An-Audio 2 introduces temporal parsing, dual text encoders, a transformer-based denoiser focused on the temporal dimension, and LLM-powered data augmentation to achieve state-of-the-art results in both objective and subjective metrics on standard T2A benchmarks (Huang et al., 2023).

1. Latent Diffusion Framework for T2A Generation

Make-An-Audio 2 models audio synthesis as a latent-space denoising diffusion process over mel-spectrogram representations. Given an input mel-spectrogram x∈RCa×Tx \in \mathbb{R}^{C_a \times T}, a 1D-convolutional variational autoencoder (VAE) encoder EE compresses xx into a latent code z=E(x)∈Rd×Lz = E(x) \in \mathbb{R}^{d \times L}, where LL reflects audio duration.

During training, diffusion operates in this latent space over TT discrete steps, with a fixed noise schedule {βt}⊂(0,1)\{\beta_t\} \subset (0,1):

  • Forward (Noising) Process: At step tt,

q(zt∣zt−1)=N(1−βtzt−1, βtI)q(z_t \mid z_{t-1}) = \mathcal{N}\left(\sqrt{1-\beta_t} z_{t-1},\, \beta_t I\right)

By induction, q(zt∣z0)=N(αˉtz0, (1−αˉt)I)q(z_t \mid z_0) = \mathcal{N}\left(\sqrt{\bar{\alpha}_t} z_0,\, (1-\bar{\alpha}_t) I\right) with EE0.

  • Reverse (Denoising) Process: A neural network EE1 parameterizes the denoising step as

EE2

where typically

EE3

and EE4.

  • Training Objective: The loss is a mean squared error in noise space:

EE5

with EE6, EE7.

EE8

with EE9 as optimal guidance scale.

This architecture prioritizes efficient learning in the more tractable latent space and allows variable-length audio generation (Huang et al., 2023).

2. Temporal Transformer-Based Diffusion Denoiser

Make-An-Audio 2 replaces the 2D U-Net backbone of earlier methods with a feed-forward transformer acting solely along the temporal axis of the latent variable xx0. This temporal modeling design includes:

The model stacks xx6 FFT blocks with xx7 heads and hidden dimensionality xx8. Attention cost scales as xx9, supporting variable-length audio inputs. This structure improves handling of true temporal relationships, avoiding artifacts from 2D “image-like” architectures that ignore audio’s sequential structure (Huang et al., 2023).

3. Temporal Parsing and Dual Text Encoder Fusion

Key to Make-An-Audio 2’s improvements in semantic and temporal alignment is temporal parsing with a dual text encoder system. Natural-language captions z=E(x)∈Rd×Lz = E(x) \in \mathbb{R}^{d \times L}0 are transformed via an LLM (e.g., GPT-4) to produce structured representations:

z=E(x)∈Rd×Lz = E(x) \in \mathbb{R}^{d \times L}1

This enables explicit modeling of event order (e.g., z=E(x)∈Rd×Lz = E(x) \in \mathbb{R}^{d \times L}2car door slammingz=E(x)∈Rd×Lz = E(x) \in \mathbb{R}^{d \times L}3paper_contentz=E(x)∈Rd×Lz = E(x) \in \mathbb{R}^{d \times L}4startz=E(x)∈Rd×Lz = E(x) \in \mathbb{R}^{d \times L}5@z=E(x)∈Rd×Lz = E(x) \in \mathbb{R}^{d \times L}6footstepsz=E(x)∈Rd×Lz = E(x) \in \mathbb{R}^{d \times L}7paper_contentz=E(x)∈Rd×Lz = E(x) \in \mathbb{R}^{d \times L}8midz=E(x)∈Rd×Lz = E(x) \in \mathbb{R}^{d \times L}9...). Extraction is performed by prompting the LLM with a fixed template.

This structured text LL0 is encoded by a fine-tuned T5 “temporal encoder” LL1. The original textual caption LL2 is also encoded via a frozen CLAP model LL3. The two embeddings are fused:

LL4

with LL5 a learned linear layer and LL6 denoting vector concatenation. This dual-encoder scheme assigns temporal-reasoning to the LLM+T5 stack while CLAP provides fidelity to semantic style and detail, significantly enhancing alignment for complex temporally-structured prompts (Huang et al., 2023).

4. Data Augmentation via LLM-Guided Synthesis

To address data scarcity for temporally-annotated T2A supervision, Make-An-Audio 2 synthesizes approximately 61k extra training samples using LLM-guided augmentation:

  1. From a database LL7 of single-label audio clips LL8, LL9 samples are chosen and temporally concatenated (with optional overlap).
  2. Each event is labeled by temporal position (“start”, “mid”, “end”, “all”).
  3. The resultant structured caption TT0 is generated.
  4. An LLM rewrites TT1 into coherent natural-language caption TT2.

Pseudocode for this augmentation:

q(zt∣z0)=N(αˉtz0, (1−αˉt)I)q(z_t \mid z_0) = \mathcal{N}\left(\sqrt{\bar{\alpha}_t} z_0,\, (1-\bar{\alpha}_t) I\right)0

This strategy yields a large, high-quality, temporally-diverse training set supporting robust learning for variable-length conditional generations (Huang et al., 2023).

5. Training Protocol and Loss Functions

The training corpus comprises 0.92M audio-text pairs (≈3.7k hours), pooled from AudioCaps, WavCaps, AudioSet, ESC-50, FSD50K, TUT, EpidemicSound, AdobeStock, UrbanSound, etc., augmented by the 61k LLM-built pairs and pseudo-prompts.

  • VAE Training: Loss combines
    • TT3 reconstruction: TT4
    • GAN realism penalty: TT5
    • Latent KL-penalty: TT6

TT7

  • Diffusion Training: Mean-squared noise regression loss TT8, using either learnable schedules TT9 or cosine schedules, and classifier-free guidance with {βt}⊂(0,1)\{\beta_t\} \subset (0,1)0.
  • Text Encoder Training: CLAP is frozen, T5 is fine-tuned on structured captions, {βt}⊂(0,1)\{\beta_t\} \subset (0,1)1 is trained end-to-end.
  • Optimization: AdamW optimizer, {βt}⊂(0,1)\{\beta_t\} \subset (0,1)2 for diffusion, batch size {βt}⊂(0,1)\{\beta_t\} \subset (0,1)3 GPUs, total {βt}⊂(0,1)\{\beta_t\} \subset (0,1)4M steps. VAE trained for {βt}⊂(0,1)\{\beta_t\} \subset (0,1)5k steps ({βt}⊂(0,1)\{\beta_t\} \subset (0,1)6).

A table summarizing the training configuration:

Component Architecture Key Hyperparameters
Denoiser FFT transformer 8 blocks, 8 heads, {βt}⊂(0,1)\{\beta_t\} \subset (0,1)7
Conv1D 1D kernel size=7, padding=3
VAE optimizer AdamW lr={βt}⊂(0,1)\{\beta_t\} \subset (0,1)8
Diffusion opt. AdamW lr={βt}⊂(0,1)\{\beta_t\} \subset (0,1)9
GPUs - tt0

6. Experimental Evaluation and Comparative Results

Make-An-Audio 2 is evaluated on AudioCaps (test) and in zero-shot settings on Clotho and AudioCaps-short. Metrics include Frechet Distance (FD), Inception Score (IS), KL divergence (KL), Frechet Audio Distance (FAD), CLAP-score, and Mean Opinion Scores for quality/faithfulness (MOS-Q/MOS-F).

  • On AudioCaps (100 DDIM steps, 937M parameters): tt1, tt2, tt3, tt4, tt5, tt6-tt7, tt8-tt9.
  • By comparison, previous SOTA (TANGO, 1.21B params): q(zt∣zt−1)=N(1−βtzt−1, βtI)q(z_t \mid z_{t-1}) = \mathcal{N}\left(\sqrt{1-\beta_t} z_{t-1},\, \beta_t I\right)0, q(zt∣zt−1)=N(1−βtzt−1, βtI)q(z_t \mid z_{t-1}) = \mathcal{N}\left(\sqrt{1-\beta_t} z_{t-1},\, \beta_t I\right)1, q(zt∣zt−1)=N(1−βtzt−1, βtI)q(z_t \mid z_{t-1}) = \mathcal{N}\left(\sqrt{1-\beta_t} z_{t-1},\, \beta_t I\right)2, q(zt∣zt−1)=N(1−βtzt−1, βtI)q(z_t \mid z_{t-1}) = \mathcal{N}\left(\sqrt{1-\beta_t} z_{t-1},\, \beta_t I\right)3, q(zt∣zt−1)=N(1−βtzt−1, βtI)q(z_t \mid z_{t-1}) = \mathcal{N}\left(\sqrt{1-\beta_t} z_{t-1},\, \beta_t I\right)4, q(zt∣zt−1)=N(1−βtzt−1, βtI)q(z_t \mid z_{t-1}) = \mathcal{N}\left(\sqrt{1-\beta_t} z_{t-1},\, \beta_t I\right)5-q(zt∣zt−1)=N(1−βtzt−1, βtI)q(z_t \mid z_{t-1}) = \mathcal{N}\left(\sqrt{1-\beta_t} z_{t-1},\, \beta_t I\right)6.
  • Predecessor Make-An-Audio: q(zt∣zt−1)=N(1−βtzt−1, βtI)q(z_t \mid z_{t-1}) = \mathcal{N}\left(\sqrt{1-\beta_t} z_{t-1},\, \beta_t I\right)7, q(zt∣zt−1)=N(1−βtzt−1, βtI)q(z_t \mid z_{t-1}) = \mathcal{N}\left(\sqrt{1-\beta_t} z_{t-1},\, \beta_t I\right)8-q(zt∣zt−1)=N(1−βtzt−1, βtI)q(z_t \mid z_{t-1}) = \mathcal{N}\left(\sqrt{1-\beta_t} z_{t-1},\, \beta_t I\right)9.

Zero-shot results on Clotho and short AudioCaps exhibit strong gains over all baselines, particularly in temporal and semantic metrics as well as on variable-length audio synthesis.

Ablations show:

  • Addition of structured T5 to CLAP yields substantial FD/IS improvements.
  • Dual-encoder fusion outperforms single T5 on all alignment criteria.
  • Replacement of 2D-U-Net with FFT backbone restores and exceeds performance for variable-length and temporally-complex data.

In summary, the innovations of Make-An-Audio 2—temporal parsing, dual-encoder fusion, temporal transformer-based diffusion, and augmented training data—are collectively responsible for advances in T2A generation quality and flexibility, establishing new state-of-the-art results for the field (Huang et al., 2023).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Make-An-Audio 2.