Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
102 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

MPE4G: Multimodal Pretrained Encoder for Co-Speech Gesture Generation (2305.15740v1)

Published 25 May 2023 in cs.CV and cs.AI

Abstract: When virtual agents interact with humans, gestures are crucial to delivering their intentions with speech. Previous multimodal co-speech gesture generation models required encoded features of all modalities to generate gestures. If some input modalities are removed or contain noise, the model may not generate the gestures properly. To acquire robust and generalized encodings, we propose a novel framework with a multimodal pre-trained encoder for co-speech gesture generation. In the proposed method, the multi-head-attention-based encoder is trained with self-supervised learning to contain the information on each modality. Moreover, we collect full-body gestures that consist of 3D joint rotations to improve visualization and apply gestures to the extensible body model. Through the series of experiments and human evaluation, the proposed method renders realistic co-speech gestures not only when all input modalities are given but also when the input modalities are missing or noisy.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Gwantae Kim (8 papers)
  2. Seonghyeok Noh (1 paper)
  3. Insung Ham (1 paper)
  4. Hanseok Ko (38 papers)
Citations (4)

Summary

We haven't generated a summary for this paper yet.