Papers
Topics
Authors
Recent
Search
2000 character limit reached

Diverse-T2M: Diverse 3D Motion Generation

Updated 9 July 2026
  • Diverse-T2M is a text-to-motion generation method that synthesizes diverse 3D human motions from text by modeling aleatoric uncertainty.
  • It employs a two-stage architecture combining VQ-VAE and a masked Transformer with a latent-space sampler to balance high text fidelity and output diversity.
  • The approach uses dual conditioning with noise and text features, enabling a tunable fusion weight that controls the trade-off between semantic consistency and diverse motion outputs.

Searching arXiv for the specified paper and closely related text-to-motion baselines to ground the article. Diverse-T2M is a text-to-motion generation method introduced in “Embracing Aleatoric Uncertainty: Generating Diverse 3D Human Motion” (Qin et al., 28 Aug 2025). It is designed for generating 3D human motions from text under two simultaneous requirements identified as central to the task: ensuring text-motion consistency and achieving generation diversity. The method is presented as a simple yet effective text-to-motion generation method that introduces uncertainty into the generation process, enabling the generation of highly diverse motions while preserving the semantic consistency of the text. Its core proposal is to use noise signals as carriers of diversity information in transformer-based methods, to construct a continuous latent space for text rather than a rigid one-to-one mapping, and to integrate a latent space sampler that introduces stochastic sampling into generation (Qin et al., 28 Aug 2025).

1. Problem Setting and Conceptual Framing

Text-to-motion generation seeks to synthesize 3D human motion sequences from natural-language descriptions. The formulation emphasized for Diverse-T2M is that recent methods can generate precise and high-quality human motions from text, but achieving diversity in the generated motions remains a significant challenge (Qin et al., 28 Aug 2025).

Diverse-T2M addresses this challenge by explicitly modeling aleatoric uncertainty. In the method description, uncertainty is not treated as an incidental by-product of stochastic decoding, but as a modeled component of the generation process. Specifically, the paper proposes a novel perspective that utilizes noise signals as carriers of diversity information in transformer-based methods, facilitating an explicit modeling of uncertainty (Qin et al., 28 Aug 2025).

The method is situated within a benchmark-driven setting using HumanML3D and KIT-ML. The stated empirical objective is not only to improve diversity metrics, but to do so while maintaining state-of-the-art performance in text consistency. This positioning matters because, within the reported framework, diversity and consistency are not treated as mutually exclusive in principle, although a guidance-dependent trade-off is also explicitly reported.

2. Two-Stage Architecture

Diverse-T2M follows the standard two-stage “discrete-then-predict” paradigm (Qin et al., 28 Aug 2025). In Stage 1, a VQ-VAE learns a codebook CC and an encoder/decoder pair (E,D)(\mathcal{E},\mathcal{D}) that map a continuous motion sequence MM into a sequence of discrete codes {ci}\{c_i\} and back. In Stage 2, a variational text-to-codes predictor takes as input a text feature gsg_s and a noise feature gϵg_\epsilon, processes each with a masked Transformer augmented with a latent-space sampler, and predicts a distribution over codebook entries at each time step. A final fusion step, described as akin to classifier-free guidance, combines the text-conditioned and noise-conditioned code distributions to produce diverse code sequences, which are then decoded by D\mathcal{D} back to 3D motion (Qin et al., 28 Aug 2025).

Three transformer-based modifications are specified. The first is a masked Transformer decoder for “next-code” prediction. The second is a latent-space sampler inserted between the Transformer’s final token embeddings and the code-probability MLP. The third is dual conditioning, text versus noise, with a learnable fusion weight ww (Qin et al., 28 Aug 2025).

This organization makes the architecture legible in terms already common in discrete generative motion modeling: motion is first quantized into code sequences, and a predictor then models the code dynamics under conditioning signals. The distinctive element is that both the conditioning pathway and the token-prediction pathway are modified to expose stochasticity as an explicit source of diversity rather than relying on a deterministic text representation alone.

3. Aleatoric Uncertainty and Dual Conditioning

To explicitly model generation-time uncertainty, Diverse-T2M conditions not only on the text feature gsg_s but also on a random noise feature gϵg_\epsilon (Qin et al., 28 Aug 2025). The noise feature is sampled as

(E,D)(\mathcal{E},\mathcal{D})0

where (E,D)(\mathcal{E},\mathcal{D})1 is the same dimension as the text embedding. During training, the method randomly drops the text or the noise with probability (E,D)(\mathcal{E},\mathcal{D})2 in a classifier-free scheme; at inference, it always uses both and fuses their predicted code distributions (Qin et al., 28 Aug 2025).

The Transformer-sampler block produces two distributions over code indices at time (E,D)(\mathcal{E},\mathcal{D})3: (E,D)(\mathcal{E},\mathcal{D})4 where (E,D)(\mathcal{E},\mathcal{D})5 is the codebook size. These are fused with a weight (E,D)(\mathcal{E},\mathcal{D})6: (E,D)(\mathcal{E},\mathcal{D})7 The paper states that larger (E,D)(\mathcal{E},\mathcal{D})8 injects more noise-driven diversity at the cost of text-fidelity (Qin et al., 28 Aug 2025).

This dual-conditioning mechanism defines the practical interpretation of aleatoric uncertainty in Diverse-T2M. The noise branch is not auxiliary regularization; it is a generation-time source of variation whose influence is directly tunable through (E,D)(\mathcal{E},\mathcal{D})9. A plausible implication is that the model separates semantically anchored structure, supplied by MM0, from diversity-inducing perturbation, supplied by MM1, while preserving a unified code prediction interface.

4. Continuous Latent Space and Sampler Design

Rather than deterministically mapping each text MM2 to a single vector, Diverse-T2M introduces a continuous latent variable MM3 that is stochastically sampled from a learned posterior (Qin et al., 28 Aug 2025). A text encoder, given as “e.g. CLIP-text,” produces

MM4

In the variational predictor, MM5 or MM6 is concatenated with masked code-token embeddings and passed through MM7 Transformer layers to obtain a fused feature MM8. An MLP posterior encoder then predicts

MM9

thereby defining a continuous latent space over {ci}\{c_i\}0 (Qin et al., 28 Aug 2025).

The latent-space sampler uses the reparameterization trick: {ci}\{c_i\}1 The sampled latent is then fed through a small MLP head with residual connections to produce logits over the {ci}\{c_i\}2 codes at each time step, followed by softmax to obtain {ci}\{c_i\}3. At inference, masked input tokens are replaced by a special {ci}\{c_i\}4 token and a fresh {ci}\{c_i\}5 is sampled for each new sequence, thereby inducing stochasticity (Qin et al., 28 Aug 2025).

The significance of this construction lies in the explicit rejection of a rigid one-to-one text representation. The paper’s description frames the latent space as continuous and stochastic rather than deterministic. This suggests that semantically valid motion realizations for a single text prompt are modeled as a distribution in latent space rather than as a single canonical trajectory.

5. Optimization Objective and Inference Procedure

Stage 1 combines VQ-VAE with residual quantization. The reported objective is

{ci}\{c_i\}6

where {ci}\{c_i\}7 is the number of quantization layers, {ci}\{c_i\}8 the {ci}\{c_i\}9th-layer residuals, and gsg_s0 stop-gradient (Qin et al., 28 Aug 2025).

Stage 2 uses a Transformer with a latent-space variational loss. The code-prediction term is

gsg_s1

that is, cross-entropy on masked code tokens. The KL term is

gsg_s2

The final loss is

gsg_s3

with gsg_s4, given for example as gsg_s5, weighting the KL term (Qin et al., 28 Aug 2025).

The inference procedure is also specified. Given text gsg_s6, the method computes gsg_s7, samples noise gsg_s8, feeds gsg_s9 through the Transformer+Sampler to obtain gϵg_\epsilon0, feeds gϵg_\epsilon1 through the same module to obtain gϵg_\epsilon2, computes

gϵg_\epsilon3

samples code index gϵg_\epsilon4 from gϵg_\epsilon5, reconstructs motion features gϵg_\epsilon6, and outputs gϵg_\epsilon7 (Qin et al., 28 Aug 2025).

A common misconception in diverse generative modeling is that diversity arises solely from sampling at the output layer. The specification of Diverse-T2M does not support that simplification. Its stochasticity is distributed across noise conditioning, continuous latent-variable sampling, and categorical code sampling, all integrated into the discrete motion-token pipeline.

6. Quantitative Results and Ablation Evidence

The reported benchmark results are summarized below exactly as presented for HumanML3D and KIT-ML (Qin et al., 28 Aug 2025).

Dataset / Method Metrics
HumanML3D — T2M-GPT MModality gϵg_\epsilon8, gϵg_\epsilon9 D\mathcal{D}0, FID D\mathcal{D}1, MM-Dist D\mathcal{D}2
HumanML3D — MoMask MModality D\mathcal{D}3, D\mathcal{D}4 D\mathcal{D}5, FID D\mathcal{D}6, MM-Dist D\mathcal{D}7
HumanML3D — Ours MModality D\mathcal{D}8, D\mathcal{D}9 ww0, FID ww1, MM-Dist ww2
KIT-ML — T2M-GPT MModality ww3, ww4 ww5, FID ww6, MM-Dist ww7
KIT-ML — MoMask MModality ww8, ww9 gsg_s0, FID gsg_s1, MM-Dist gsg_s2
KIT-ML — Ours MModality gsg_s3, gsg_s4 gsg_s5, FID gsg_s6, MM-Dist gsg_s7

The metric definitions are also stated explicitly. MultiModality measures average pairwise feature-distance among 30 generations per text, so higher values indicate more diverse generations. R-Precision is Top-1 motion-to-text retrieval accuracy. FID and MM-Dist gauge realism and text-motion alignment, with lower values preferred (Qin et al., 28 Aug 2025).

The ablation on HumanML3D isolates the method’s components. The baseline without latent sampler, reported as “VQ+Transformer,” gives MModality gsg_s8 and gsg_s9. Adding the Variational Predictor only gives MModality gϵg_\epsilon0 with gϵg_\epsilon1. Adding Noise Conditioning only gives MModality gϵg_\epsilon2 with gϵg_\epsilon3. The full model, denoted “RVQ+VP+Noise,” gives MModality gϵg_\epsilon4 and gϵg_\epsilon5 (Qin et al., 28 Aug 2025).

These ablations matter because they separate two distinct contributors to diversity. The variational predictor improves diversity relative to the baseline, while noise conditioning produces a substantially larger increase. The full model then combines both while improving gϵg_\epsilon6. This suggests that the reported diversity gains are not attributable to a single stochastic component alone.

7. Diversity–Consistency Trade-off and Qualitative Behavior

Diverse-T2M explicitly presents balancing diversity and consistency as a controllable property rather than a fixed operating point. The guidance weight sweep over gϵg_\epsilon7 shows that increasing gϵg_\epsilon8 raises diversity, measured by MModality, but slowly degrades R-Precision and MM-Dist, illustrating a smooth trade-off curve (Qin et al., 28 Aug 2025).

This trade-off is consistent with the fusion rule

gϵg_\epsilon9

under which larger (E,D)(\mathcal{E},\mathcal{D})00 increases the influence of the noise-conditioned distribution. The paper also states this directly: larger (E,D)(\mathcal{E},\mathcal{D})01 injects more noise-driven diversity at the cost of text-fidelity (Qin et al., 28 Aug 2025). A plausible implication is that Diverse-T2M is intended not merely as a single model, but as a tunable generation framework whose operating regime can be selected according to whether a downstream setting prioritizes exploration or strict semantic adherence.

The qualitative examples reported for the method show that it can vary “walking direction,” “which hand picks up an object,” or “wiping posture,” while still faithfully following multi-step instructions such as “pick–walk–pick–wipe” (Qin et al., 28 Aug 2025). Within the terms of the paper, these examples are important because they illustrate that diversity is expressed through semantically compatible alternatives rather than arbitrary motion perturbation.

A second common misconception is that diversity necessarily entails loss of instruction fidelity across complex prompts. The reported qualitative examples and the benchmark (E,D)(\mathcal{E},\mathcal{D})02 values do not support such a blanket conclusion. At the same time, the guidance-weight sweep documents that the trade-off is real and gradual rather than absent. The paper therefore frames diversity and text consistency as jointly optimizable but not completely decoupled.

8. Position Within Text-to-Motion Modeling

Diverse-T2M is presented as a transformer-based method that remains within the standard two-stage “discrete-then-predict” paradigm while altering the semantics of conditioning and decoding through explicit uncertainty modeling (Qin et al., 28 Aug 2025). Relative to deterministic text-to-code mappings, its principal distinction is the construction of a continuous latent space and the use of stochastic sampling in that space. Relative to purely text-conditioned predictors, its principal distinction is the introduction of a random noise feature as an explicit conditioning input.

The reported results on HumanML3D and KIT-ML support the paper’s central claim that the method significantly enhances diversity while maintaining state-of-the-art performance in text consistency (Qin et al., 28 Aug 2025). In summary form provided by the paper itself, the combination of stochastic sampling in a learned latent space and noise-conditioning with a guiding fusion weight injects explicit aleatoric uncertainty, producing much richer and more realistic variations of 3D motion for a given textual command without sacrificing text-conditioned accuracy (Qin et al., 28 Aug 2025).

Within the landscape suggested by the paper’s own comparisons, Diverse-T2M can therefore be understood as a discrete text-to-motion generator in which uncertainty is promoted from an implicit artifact of sampling to an architectural and probabilistic design principle.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Diverse-T2M.