MACS: Mass Conditioned 3D Hand and Object Motion Synthesis
Abstract: The physical properties of an object, such as mass, significantly affect how we manipulate it with our hands. Surprisingly, this aspect has so far been neglected in prior work on 3D motion synthesis. To improve the naturalness of the synthesized 3D hand object motions, this work proposes MACS the first MAss Conditioned 3D hand and object motion Synthesis approach. Our approach is based on cascaded diffusion models and generates interactions that plausibly adjust based on the object mass and interaction type. MACS also accepts a manually drawn 3D object trajectory as input and synthesizes the natural 3D hand motions conditioned by the object mass. This flexibility enables MACS to be used for various downstream applications, such as generating synthetic training data for ML tasks, fast animation of hands for graphics workflows, and generating character interactions for computer games. We show experimentally that a small-scale dataset is sufficient for MACS to reasonably generalize across interpolated and extrapolated object masses unseen during the training. Furthermore, MACS shows moderate generalization to unseen objects, thanks to the mass-conditioned contact labels generated by our surface contact synthesis model ConNet. Our comprehensive user study confirms that the synthesized 3D hand-object interactions are highly plausible and realistic.
- TensorFlow: Large-scale machine learning on heterogeneous systems, 2015. Software available from tensorflow.org.
- G. Bradski. The OpenCV Library. Dr. Dobb’s Journal of Software Tools, 2000.
- Contactdb: Analyzing and predicting grasp contact via thermal imaging. In Computer Vision and Pattern Recognition (CVPR), 2019.
- D-grasp: Physically plausible dynamic grasp synthesis for hand-object interactions. In Computer Vision and Pattern Recognition (CVPR), 2022.
- Fast and accurate deep network learning by exponential linear units (elus). arXiv preprint arXiv:1511.07289, 2015.
- Ganhand: Predicting human grasp affordances in multi-object scenes. In Computer Vision and Pattern Recognition (CVPR), 2020.
- Mofusion: A framework for denoising-diffusion-based motion synthesis. In Computer Vision and Pattern Recognition (CVPR), 2023.
- Refining grasp affordance models by experience. In International Conference on Robotics and Automation (ICRA), 2010.
- Imos: Intent-driven full-body motion synthesis for human-object interactions. In Eurographics, 2023.
- Generative adversarial nets. In Advances in Neural Information Processing Systems (NeurIPS), 2014.
- Contactopt: Optimizing contact to improve grasps. In Computer Vision and Pattern Recognition (CVPR), 2021.
- Action2motion: Conditioned generation of 3d human motions. In ACM International Conference on Multimedia, 2020.
- Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems (NeurIPS), 2020.
- Physical interaction: Reconstructing hand-object interactions with physics. In SIGGRAPH Asia 2022 Conference Papers, 2022.
- Grasping field: Learning implicit representations for human grasps. In International Conference on 3D Vision (3DV), 2020.
- Auto-encoding variational bayes. In International Conference on Learning Representations (ICLR), 2014.
- Diffwave: A versatile diffusion model for audio synthesis. In International Conference on Learning Representations (ICLR), 2021.
- On the efficient computation of independent contact regions for force closure grasps. In International Conference on Intelligent Robots and Systems (ICIRS), 2010.
- Data-driven grasp synthesis using shape matching and task-based pruning. Transactions on visualization and computer graphics (TVCG), 13(4):732–747, 2007.
- Semi-supervised 3d hand-object poses estimation with interactions in time. In Computer Vision and Pattern Recognition (CVPR), 2021.
- Mediapipe: A framework for perceiving and processing reality. In Workshop on Computer Vision for AR/VR at Computer Vision and Pattern Recognition (CVPRW), 2019.
- Contact-invariant optimization for hand manipulation. In Proceedings of the ACM SIGGRAPH/Eurographics symposium on computer animation (SCA), pages 137–144, 2012.
- Real-time pose and shape reconstruction of two interacting hands with a single depth camera. ACM Transactions on Graphics (TOG), 38(4), 2019.
- Physically based grasping control from example. In ACM SIGGRAPH/Eurographics symposium on Computer animation (SCA), 2005.
- Dreamfusion: Text-to-3d using 2d diffusion. In International Conference on Learning Representations (ICLR), 2023.
- High-resolution image synthesis with latent diffusion models. In Computer Vision and Pattern Recognition (CVPR), 2022.
- Photorealistic text-to-image diffusion models with deep language understanding. In Advances in neural information processing systems (NeurIPS), 2022.
- Hand-object interaction detection with fully convolutional networks. In Computer Vision and Pattern Recognition (CVPR), 2017.
- Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning (ICML), 2015.
- Learning structured output representation using deep conditional generative models. Advances in neural information processing systems (NeurIPS), 2015.
- GRAB: A dataset of whole-body human grasping of objects. In European Conference on Computer Vision (ECCV), 2020.
- Goal: Generating 4d whole-body motion for hand-object grasping. In Computer Vision and Pattern Recognition (CVPR), 2022.
- H+ o: Unified egocentric recognition of 3d hand-object poses and interactions. In Computer Vision and Pattern Recognition (CVPR), 2019.
- Human motion diffusion model. In International Conference on Learning Representations (ICLR), 2023.
- Imitation learning of hand gestures and its evaluation for humanoid robots. In International Conference on Information and Automation (ICIA), 2010.
- Rgb2hands: real-time tracking of 3d hand interactions from monocular rgb video. ACM Transactions on Graphics (TOG), 39(6), 2020.
- Ghum & ghuml: Generative 3d human shape and articulated pose models. In Computer Vision and Pattern Recognition (CVPR), 2020.
- Synthesis of detailed hand manipulations using contact sampling. ACM Transactions on Graphics (ToG), 31(4), 2012.
- Vaegan: A collaborative filtering framework based on adversarial variational autoencoders. In International Joint Conference on Artificial Intelligence (IJCAI), 2019.
- Physdiff: Physics-guided human motion diffusion model. In International Conference on Computer Vision (ICCV), 2023.
- Manipnet: Neural manipulation synthesis with a hand-object spatial representation. ACM Transactions on Graphics (TOG), 40(4):1–14, 2021.
- Motiondiffuse: Text-driven human motion generation with diffusion model. arXiv preprint arXiv:2208.15001, 2022.
- Cams: Canonicalized manipulation spaces for category-level functional hand-object manipulation synthesis. In Computer Vision and Pattern Recognition (CVPR), 2023.
- Toch: Spatio-temporal object correspondence to hand for motion refinement. In European Conference on Computer Vision (ECCV), 2022.
- On the continuity of rotation representations in neural networks. In Computer Vision and Pattern Recognition (CVPR), 2019.
Paper Prompts
Sign up for free to create and run prompts on this paper.