GoMAvatar: Efficient Animatable Human Modeling from Monocular Video Using Gaussians-on-Mesh
Abstract: We introduce GoMAvatar, a novel approach for real-time, memory-efficient, high-quality animatable human modeling. GoMAvatar takes as input a single monocular video to create a digital avatar capable of re-articulation in new poses and real-time rendering from novel viewpoints, while seamlessly integrating with rasterization-based graphics pipelines. Central to our method is the Gaussians-on-Mesh representation, a hybrid 3D model combining rendering quality and speed of Gaussian splatting with geometry modeling and compatibility of deformable meshes. We assess GoMAvatar on ZJU-MoCap data and various YouTube videos. GoMAvatar matches or surpasses current monocular human modeling algorithms in rendering quality and significantly outperforms them in computational efficiency (43 FPS) while being memory-efficient (3.63 MB per subject).
- Video based reconstruction of 3D people models. In CVPR, 2018.
- SCAPE: shape completion and animation of people. ACM TOG, 2005.
- Learning Neural Light Fields with Ray-Space Embedding. In CVPR, 2021.
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields. In ICCV, 2021.
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields. In CVPR, 2022.
- Zip-NeRF: Anti-Aliased Grid-Based Neural Radiance Fields. In ICCV, 2023.
- TensoRF: Tensorial Radiance Fields. In ECCV, 2022a.
- Animatable Neural Radiance Fields from Monocular RGB Video. arXiv, 2021a.
- GM-NeRF: Learning Generalizable Model-based Neural Radiance Fields from Multi-view Images. In CVPR, 2023a.
- SNARF: Differentiable Forward Skinning for Animating Non-Rigid Neural Implicit Shapes. In ICCV, 2021b.
- Fast-SNARF: A Fast Deformer for Articulated Neural Fields. TPAMI, 2022b.
- Learning Implicit Fields for Generative Shape Modeling. In CVPR, 2019.
- MobileNeRF: Exploiting the Polygon Rasterization Pipeline for Efficient Neural Field Rendering on Mobile Architectures. In CVPR, 2023b.
- MPS-NeRF: Generalizable 3D Human Rendering from Multiview Images. TPAMI, 2022.
- FastNeRF: High-Fidelity Neural Rendering at 200FPS. In ICCV, 2021.
- Learning Neural Volumetric Representations of Dynamic Humans in Minutes. In CVPR, 2023.
- The lumigraph. In SIGGRAPH, 1996.
- Shape, Light & Material Decomposition from Images using Monte Carlo Rendering and Denoising. In NeurIPS, 2022.
- Arch++: Animation-ready clothed human reconstruction revisited. In ICCV, 2021.
- Baking Neural Radiance Fields for Real-Time View Synthesis. In ICCV, 2021.
- SHERF: Generalizable Human NeRF from a Single Image. In ICCV, 2023.
- HVTR: Hybrid Volumetric-Textural Rendering for Human Avatars. 3DV, 2021.
- Let there be color! ACM TOG, 2016.
- NASA: Neural Articulated Shape Approximation. In ECCV, 2020.
- SelfRecon: Self Reconstruction Your Digital Avatar from Monocular Video. In CVPR, 2022a.
- Instantavatar: Learning avatars from monocular video in 60 seconds. In CVPR, 2023.
- Neuman: Neural human radiance field from a single video. In ECCV, 2022b.
- Panoptic Studio: A Massively Multiview System for Social Interaction Capture. TPAMI, 2017.
- Ray Tracing Volume Densities. In SIGGRAPH, 1984.
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM TOG, 2023.
- You Only Train Once: Multi-Identity Free-Viewpoint Neural Human Rendering from Monocular Videos. arXiv, 2023.
- Adam: A Method for Stochastic Optimization. arXiv, 2014.
- PARE: Part attention regressor for 3D human body estimation. In ICCV, 2021.
- Neural Human Performer: Learning Generalizable Radiance Fields for Human Performance Rendering. In NeurIPS, 2021.
- Neural Image-based Avatars: Generalizable Radiance Fields for Human Avatar Modeling. In ICLR, 2023.
- Light Field Rendering. In SIGGRAPH, 1996.
- TAVA: Template-free Animatable Volumetric Actors. In ECCV, 2022.
- Neurmips: Neural Mixture of Planar Experts for View Synthesis. In CVPR, 2022.
- Neural actor: Neural free-view synthesis of human actors with pose control. ACM TOG, 2021.
- Soft rasterizer: A differentiable renderer for image-based 3d reasoning. In ICCV, 2019.
- Neural Volumes: Learning Dynamic Renderable Volumes from Images. ACM TOG, 2019.
- Mosh: Motion and shape capture from sparse markers. ACM TOG, 2014.
- SMPL: A skinned multi-person linear model. ACM TOG, 2015.
- Mediapipe: A framework for building perception pipelines. arXiv, 2019.
- Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis. In 3DV, 2024.
- NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections. In CVPR, 2021.
- Occupancy networks: Learning 3D reconstruction in function space. In CVPR, 2019.
- COAP: Compositional Articulated Occupancy of People. In CVPR, 2022.
- NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. In ECCV, 2020.
- Instant neural graphics primitives with a multiresolution hash encoding. ACM TOG, 2022.
- DeepSDF: Learning continuous signed distance functions for shape representation. In CVPR, 2019.
- Capturing and animating skin deformation in human motion. ACM TOG, 2006.
- Animatable Neural Radiance Fields for Modeling Dynamic Human Bodies. In ICCV, 2021a.
- Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In CVPR, 2021b.
- Drivable Volumetric Avatars using Texel-Aligned Features. In SIGGRAPH, 2022.
- Class-agnostic Reconstruction of Dynamic Objects from Videos. In NeurIPS, 2021.
- Neural volumetric object selection. In CVPR, 2022.
- Real-time Volumetric Rendering of Dynamic Humans. arXiv, 2023.
- PIFu: Pixel-aligned implicit function for high-resolution clothed human digitization. In ICCV, 2019.
- PIFuHD: Multi-level pixel-aligned implicit function for high-resolution 3d human digitization. In CVPR, 2020.
- Layered depth images. In SIGGRAPH, 1998.
- 3D Photography Using Context-Aware Layered Depth Inpainting. In CVPR, 2020.
- DeepVoxels: Learning Persistent 3D Feature Embeddings. In CVPR, 2019.
- A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and Pose. In NeurIPS, 2021.
- Human motion diffusion model. In ICLR, 2023.
- Neural-GIF: Neural Generalized Implicit Functions for Animating People in Clothing. In ICCV, 2021.
- Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields. In CVPR, 2022.
- Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. In NeurIPS, 2021.
- ARAH: Animatable Volume Rendering of Articulated Human SDFs. In ECCV, 2022.
- HumanNeRF: Free-viewpoint Rendering of Moving People from Monocular Video. In CVPR, 2022.
- NeX: Real-time View Synthesis with Neural Basis Expansion. In CVPR, 2021.
- 4d gaussian splatting for real-time dynamic scene rendering. arXiv, 2023a.
- 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering. arXiv, 2023b.
- Multiface: A Dataset for Neural Face Rendering, 2022.
- H-NeRF: Neural Radiance Fields for Rendering and Temporal Reconstruction of Humans in Motion. In NeurIPS, 2021.
- Person search in a scene by jointly modeling people commonness and person uniqueness. In ACM MM, 2014.
- S3: Neural shape, skeleton, and skinning fields for 3d human modeling. In CVPR, 2021.
- Bakedsdf: Meshing neural sdfs for real-time view synthesis. In SIGGRAPH, 2023.
- PlenOctrees for Real-time Rendering of Neural Radiance Fields. In ICCV, 2021.
- Plenoxels: Radiance Fields without Neural Networks. In CVPR, 2022.
- Monohuman: Animatable human neural field from monocular video. In CVPR, 2023.
- NeRF++: Analyzing and Improving Neural Radiance Fields. arXiv, 2020.
- NDF: Neural Deformable Fields for Dynamic Human Modelling. In ECCV, 2022.
- The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018.
- HumanNeRF: Efficiently Generated Human Radiance Field from Sparse Inputs. In CVPR, 2022.
- Occupancy Planes for Single-view RGB-D Human Reconstruction. In AAAI, 2023.
- Pseudo-Generalized Dynamic View Synthesis from a Video. In ICLR, 2024.
- PointAvatar: Deformable Point-based Head Avatars from Videos. In CVPR, 2023.
- Structured Local Radiance Fields for Human Avatar Modeling. In CVPR, 2022.
- Stereo Magnification: Learning View Synthesis using Multiplane Images. ACM TOG, 2018.
Paper Prompts
Sign up for free to create and run prompts on this paper.