SMooGPT: Stylized Motion Generation
- SMooGPT is a stylized motion generation framework that reformulates motion synthesis as a reasoning-composition-generation process using a body-part text space.
- It integrates part-specific VQ-VAE tokenization with LLM-based instruction tuning to enable fine-grained control and interpretable style transfer.
- The approach achieves state-of-the-art style recognition and robust generalization, addressing common latent-space limitations in motion synthesis.
to=functions.arxiv_search 天天中彩票不能 天天中彩票官网ide the query. {"query":"(Zhong et al., 4 Sep 2025) SMooGPT Stylized Motion Generation using LLMs", "max_results": 5} to=functions.arxiv_search _影音先锋 彩票直属 手机天天彩票? {"query":"SMooGPT Stylized Motion Generation using LLMs", "max_results": 10} to=functions.arxiv_search /久久 大发快三是不是 彩神争霸快三json {"query":"SMooGPT Stylized Motion Generation using LLMs", "max_results": 10} to=functions.arxiv_search _一本道 րաստիճան 】【:】【“】【json {"query":"SMooGPT Stylized Motion Generation using LLMs", "max_results": 10} SMooGPT is a stylized motion generation framework that reformulates the task of producing a motion consistent with both motion content and motion style as a reasoning-composition-generation process in a body-part text space, rather than as direct mixing in latent motion space (Zhong et al., 4 Sep 2025). The target problem is to generate motions such as “walking in a loop like a Monkey,” while preserving content and reflecting style. Its central premise is that human motion is naturally compositional, that body-part behavior can often be described effectively in natural language, and that LLMs can exploit this structure to produce more interpretable, controllable, and open-vocabulary stylization than prior latent-space approaches (Zhong et al., 4 Sep 2025).
1. Problem formulation and motivation
Stylized motion generation in computer graphics seeks to synthesize a novel motion that respects both a specified motion content and a desired motion style. Prior work described in the SMooGPT paper mainly addresses this objective through motion style transfer or conditional motion generation, typically by embedding style in a latent vector and fusing it with content features or injecting it into a generative model (Zhong et al., 4 Sep 2025).
In the reported formulation, these latent-space approaches have three recurrent limitations. First, they are difficult to interpret because style influence is encoded implicitly. Second, they offer limited fine-grained control because content and style are combined globally rather than at the level of distinct body parts. Third, they generalize poorly to new styles, partly because curated stylization datasets such as 100STYLE exhibit a strong locomotion bias, with many styles effectively appearing as variations of walking (Zhong et al., 4 Sep 2025).
SMooGPT is motivated by three observations stated explicitly in the paper: human motion can often be described in a body-part-centric natural language form; LLMs have strong capabilities for understanding and reasoning about human motion; and human motion is inherently compositional, so new content or style can be generated by recomposing familiar local behaviors. This suggests that stylization can be treated not primarily as feature interpolation, but as structured reasoning over part-level motion descriptions.
2. Body-part text space
The core intermediate representation is the body-part text space
where denotes the motion, a global text description, and a set of part-specific texts for six body parts: Root, Backbone, Left Arm, Right Arm, Left Leg, and Right Leg (Zhong et al., 4 Sep 2025).
Within this representation, the global text carries coarse semantics and high-level style, whereas the body-part texts encode localized spatial and dynamic behavior, including direction, posture, rhythm, and motion intensity. The body-part text space is therefore not an auxiliary annotation layer but the central computational substrate of the framework: instead of composing motions directly, SMooGPT composes descriptions of how each body part should move (Zhong et al., 4 Sep 2025).
This part-wise representation is designed to make stylization interpretable. Because the generated specification is expressed at the level of individual body parts, content-style conflicts can be addressed explicitly in language. It also supports reuse: the paper argues that new styles can often be synthesized by recombining body-part behaviors already seen during training. A plausible implication is that the open-vocabulary capability of the LLM is useful not because it memorizes new style labels, but because it can decompose unseen style descriptions into familiar local motion attributes.
3. Architecture and motion tokenization
SMooGPT uses HumanML3D-style motion features, with a motion sequence
The sequence is split into six part motions , one for each body part (Zhong et al., 4 Sep 2025).
Each part is modeled by a part-specific VQ-VAE. For body part , the framework uses an encoder , a codebook
and a decoder . Quantization is defined as
0
The corresponding VQ-VAE objective is
1
with stop-gradient 2 and commitment weight 3 (Zhong et al., 4 Sep 2025).
These discrete part tokens are incorporated into the vocabulary of a pre-trained Flan-T5-Base model. The unified vocabulary is
4
where 5 is the original text vocabulary, 6 denotes body-part motion tokens, and 7 contains special delimiters such as segment boundaries for each body part. The model is trained as an encoder-decoder transformer with the autoregressive language modeling loss
8
A potential misconception is that SMooGPT is merely a text-only planner placed in front of an external motion generator. In the reported architecture, the fine-tuned LLM also generates body-part motion tokens, which are then decoded back into full-body motion (Zhong et al., 4 Sep 2025).
4. Reasoning, composition, and instruction tuning
Training proceeds in two major stages. The first is pre-training for modality alignment, in which the model learns bidirectional translation between body-part motion tokens and body-part texts, effectively treating motion as a foreign language. The purpose of this stage is to align motion and language in a shared token space and to teach the model the “grammar” of part-wise motion (Zhong et al., 4 Sep 2025).
The second is post-training with instruction following. Here the model is fine-tuned with explicit prompts for three tasks: translating a global motion text into detailed body-part texts, composing content and style body-part texts into a unified conflict-free set, and generating motion tokens from body-part texts. The dataset is curated from HumanML3D and 100STYLE, with body-part annotations largely produced using ChatGPT-3.5. HumanML3D provides content motion-text pairs, while 100STYLE contributes style motions and style labels. The supplementary additionally describes a stylization-enhanced tuning stage using content/body-part and style/body-part text pairs (Zhong et al., 4 Sep 2025).
At inference, the same three functional roles organize the workflow. As a reasoner, the model converts either a motion sequence or a text description into structured body-part descriptions. As a composer, it merges content and style descriptions into a coherent part-level specification. As a generator, it emits body-part motion tokens that are decoded into full-body motion.
The explicit composition stage is intended to resolve conflicts that are difficult to handle in latent-space fusion. The paper’s illustrative example contrasts a “throw” motion, which may require one arm close to the body, with an “Aeroplane” style, which demands extended arms. In SMooGPT, the composer can borrow compatible aspects from both descriptions and produce a consistent body-part instruction set, rather than forcing a brittle global compromise. Competing latent-space methods are described as prone to content mismatch, style collapse, or overwriting of constraints in such cases (Zhong et al., 4 Sep 2025).
5. Generalization and empirical evaluation
The evaluation covers text-guided stylization and motion-guided stylization on HumanML3D and 100STYLE. For text-guided stylization, the baselines are ChatGPT + MotionGPT and ChatGPT + MLD. For motion-guided stylization, the comparisons include SMooDi, an in-domain version, an out-of-domain version, and SMooDi9 (Zhong et al., 4 Sep 2025).
The paper organizes evaluation around three axes: content preservation, style reflection, and realism. The corresponding metrics are R-Precision (Top-3) and MM Dist for content preservation, Style Recognition Accuracy (SRA) for style reflection, and FS-Ratio for motion realism. The authors explicitly avoid using FID for content preservation because FID is not conditioning-aware and becomes ambiguous in stylized generation settings where style intentionally alters the motion distribution (Zhong et al., 4 Sep 2025).
In the pure text-driven stylized motion generation setting, SMooGPT is reported as achieving a new state of the art. Its headline style-reflection result is SRA 33.33, compared with 7.61 for ChatGPT + MotionGPT and 4.82 for ChatGPT + MLD. The paper also reports MM Dist 5.096, R-Precision 0.563, and FS-Ratio 0.141 for SMooGPT in this setting.
| Method | SRA |
|---|---|
| ChatGPT + MotionGPT | 7.61 |
| ChatGPT + MLD | 4.82 |
| SMooGPT | 33.33 |
For motion-guided stylization, the paper states that SMooDi performs better on in-domain seen styles because classifier guidance is strong for familiar categories, whereas SMooGPT is stronger on out-of-domain styles and is much less sensitive to dataset bias. The reported pattern is that SMooGPT’s style accuracy drops only modestly from in-domain to out-of-domain, while SMooDi undergoes a much larger decline. The stated explanation is that LLM-based body-part reasoning can generalize to styles such as “Marching,” “Stage Performance,” and “Cheering” even when those styles are not explicitly labeled in the training data (Zhong et al., 4 Sep 2025).
Qualitatively, the paper highlights examples including “Kungfu kick,” “Aeroplane jumping,” and “Lazy punch.” The accompanying visualizations expose the simplified body-part text used internally, making the generation pathway interpretable. The conflict-resolution case involving a throwing action and an Aeroplane pose is used to show that SMooGPT can assign different arm behaviors rather than collapse the motion into an incoherent whole.
6. Ablations, perceptual study, and limitations
The user study recruits 50 participants and asks them to compare generated motions pairwise on content preservation, style reflection, and motion realism. Participants prefer SMooGPT over ChatGPT+MotionGPT about 81% of the time across all three criteria, and over ChatGPT+MLD by more than 72%. The paper presents this perceptual evidence as aligned with the quantitative metrics (Zhong et al., 4 Sep 2025).
The ablation studies isolate two components as especially important. Removing the body-part space and reverting to a global-text-only MotionGPT-style setup substantially reduces style recognition, with SRA dropping from 33.33 to 17.96. This is presented as evidence that the structured part-wise representation is crucial for fine-grained control. Removing pre-training also degrades all metrics, especially SRA, indicating that the motion-text alignment stage is necessary before instruction tuning. The supplementary further reports that pretraining/fine-tuning on HumanML3D improves FID, MM Dist, R-Precision, and FS-Ratio modestly, suggesting continued refinement of motion generation quality after the instruction-following phases (Zhong et al., 4 Sep 2025).
The principal strengths identified in the paper are interpretability, fine-grained controllability, and generalization to new styles through open-vocabulary language reasoning. The main limitation explicitly acknowledged by the authors is that the current body-part text space remains coarse in temporal expressiveness: it captures localized motion reasonably well but is not yet rich enough to fully encode intricate temporal dynamics. The proposed future direction is to extend each part description with keypose-based temporal text (Zhong et al., 4 Sep 2025).
Taken together, SMooGPT’s contribution is not only a specific model architecture but a reformulation of stylized motion synthesis as a language-driven, body-part-wise reasoning and composition problem. This suggests a broader shift in how controllable motion generation can be organized: from implicit latent blending toward explicit intermediate descriptions whose semantics are inspectable, editable, and recombinable.