Achieving fine-grained spatiotemporal control in human motion generation
Develop human motion generation models that achieve fine-grained simultaneous control over spatial structure at the per-body-part level and temporal dynamics across motion sequences, enabling precise alignment of generated motions with detailed spatiotemporal constraints.
References
Despite these advances, achieving fine-grained spatial and temporal control in motion generation remains a challenging open problem.
— FrankenMotion: Part-level Human Motion Generation and Composition
(2601.10909 - Li et al., 15 Jan 2026) in Section 2 (Related Work), Motion generation with control
It would be interesting to investigate whether a two-stage approach could further refine the generated motion by dividing the task: for example, first predicting stick-tip positions from audio and subsequently synthesizing the final motion based on these targets.
— Generalized Audio-Driven Synthesis of Precise Drummer Motion
(2608.19055 - Iñesta et al., 19 Aug 2026) in Section 6, Discussion, paragraph “Future directions”