Towards Open Domain Text-Driven Synthesis of Multi-Person Motions (2405.18483v2)

Published 28 May 2024 in cs.CV

Abstract: This work aims to generate natural and diverse group motions of multiple humans from textual descriptions. While single-person text-to-motion generation is extensively studied, it remains challenging to synthesize motions for more than one or two subjects from in-the-wild prompts, mainly due to the lack of available datasets. In this work, we curate human pose and motion datasets by estimating pose information from large-scale image and video datasets. Our models use a transformer-based diffusion framework that accommodates multiple datasets with any number of subjects or frames. Experiments explore both generation of multi-person static poses and generation of multi-person motion sequences. To our knowledge, our method is the first to generate multi-subject motion sequences with high diversity and fidelity from a large variety of textual prompts.

Authors (8)

Mengyi Shan (10 papers)
Lu Dong (17 papers)
Yutao Han (5 papers)
Yuan Yao (292 papers)
Tao Liu (350 papers)
Ifeoma Nwogu (18 papers)
Guo-Jun Qi (76 papers)
Mitch Hill (9 papers)

Citations (4)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Tweets

https://twitter.com/gastronomy/status/1796031037098586589

https://twitter.com/CSVisionPapers/status/1796034167156637835

Towards Open Domain Text-Driven Synthesis of Multi-Person Motions (2405.18483v2)

Summary

Related Papers

Tweets