Papers
Topics
Authors
Recent
Search
2000 character limit reached

Intrinsic Geometry-Appearance Consistency Optimization for Sparse-View Gaussian Splatting

Published 3 Mar 2026 in cs.CV | (2603.02893v1)

Abstract: 3D human reconstruction from a single image is a challenging problem and has been exclusively studied in the literature. Recently, some methods have resorted to diffusion models for guidance, optimizing a 3D representation via Score Distillation Sampling(SDS) or generating a back-view image for facilitating reconstruction. However, these methods tend to produce unsatisfactory artifacts (\textit{e.g.} flattened human structure or over-smoothing results caused by inconsistent priors from multiple views) and struggle with real-world generalization in the wild. In this work, we present \emph{MVD-HuGaS}, enabling free-view 3D human rendering from a single image via a multi-view human diffusion model. We first generate multi-view images from the single reference image with an enhanced multi-view diffusion model, which is well fine-tuned on high-quality 3D human datasets to incorporate 3D geometry priors and human structure priors. To infer accurate camera poses from the sparse generated multi-view images for reconstruction, an alignment module is introduced to facilitate joint optimization of 3D Gaussians and camera poses. Furthermore, we propose a depth-based Facial Distortion Mitigation module to refine the generated facial regions, thereby improving the overall fidelity of the reconstruction. Finally, leveraging the refined multi-view images, along with their accurate camera poses, MVD-HuGaS optimizes the 3D Gaussians of the target human for high-fidelity free-view renderings. Extensive experiments on Thuman2.0 and 2K2K datasets show that the proposed MVD-HuGaS achieves state-of-the-art performance on single-view 3D human rendering.

Summary

  • The paper introduces ICO-GS, a sparse-view 3D Gaussian Splatting method that jointly regularizes geometry and appearance through feature-based multi-view consistency, edge-aware depth smoothing, and filtered virtual-view supervision.
  • ICO-GS achieves 22.20 dB PSNR on three-view LLFF and 21.77 dB on three-view DTU, improving over the strongest prior methods by 0.76 dB and 1.06 dB, respectively.
  • The method improves weakly textured regions and depth boundaries without external monocular depth priors, but increases training time by about 1.5× and remains vulnerable to specular or view-dependent appearance.

The geometry-appearance discrepancy in sparse-view 3DGS

3D Gaussian Splatting (3DGS) parameterizes each primitive with coupled intrinsic properties: geometric attributes (position μ\boldsymbol{\mu}, covariance Σ\boldsymbol{\Sigma}, opacity α\alpha) and appearance attributes (view-dependent color). The paper argues that faithful sparse-view reconstruction requires intrinsic consistency between these attribute groups—geometry must capture true 3D structure while appearance reflects coherent surface photometry—and that standard 3DGS optimization violates this under sparse observations. The authors demonstrate a concrete failure mode: as training views decrease from dense to 2–3 inputs, training-view RGB remains well-fitted while rendered depth collapses into noise and floaters. Two mechanisms are identified. First, the photometric loss constrains only projected appearance, not depth, so Gaussians can slide freely along camera rays when multi-view overlap is insufficient. Second, an "appearance compensation" shortcut lets the optimizer adjust color and opacity to mask misplaced Gaussians rather than correcting their positions, producing high training-view PSNR despite incorrect geometry.

This diagnosis positions the paper against two families of prior remedies: monocular depth priors (FSGS, DNGaussian), which introduce scale ambiguity and estimator noise, and rendered-depth-based virtual view supervision (BinocularGS), which risks a circular dependency where unreliable depth yields misaligned virtual views that further corrupt geometry. ICO-GS is designed to break this loop by regularizing geometry first and filtering depth reliability before it is used for appearance supervision.

Robust geometric regularization

The geometric component warps source views to a reference view using alpha-blended rendered depth and penalizes inconsistency. Three design choices address known failure modes of naive photometric matching:

  • Feature-based matching: raw RGB consistency is replaced by cosine distance between features from a frozen pre-trained MVS network (from CREStereo/cascade cost volume work), computed once during preprocessing, which tolerates illumination variation and specular effects at negligible runtime cost.
  • Pixel-wise top-kk selection: for each reference pixel, only the k=⌈(n−1)/2⌉k = \lceil(n-1)/2\rceil most consistent source correspondences contribute to the loss, so pixels occluded in some views still receive valid supervision from the remainder.
  • Edge-aware depth smoothness: a gradient-weighted total-variation term on depth, modulated by image gradients (α=1\alpha=1), regularizes monocularly visible regions while preserving depth discontinuities at object boundaries.

The paper claims these constraints are particularly effective in weakly-textured regions, where photometric cues alone cannot disambiguate geometry; the DTU results below support this claim quantitatively.

Geometry-guided appearance optimization

Rather than trusting all rendered depth for virtual view synthesis, the method validates depth through cycle-consistency filtering (CCDF): each pixel is forward-warped to a source view using reference depth, then back-warped using source depth. A pixel is deemed reliable if its forward-backward depth error falls below τd=0.01⋅max⁡(D0)\tau_d = 0.01 \cdot \max(D_0) in at least m=⌈(n−1)/2⌉m = \lceil(n-1)/2\rceil source views. Virtual poses are sampled within a sphere of radius rr around each reference position—broader than the fixed binocular pairs of BinocularGS—and virtual images are synthesized by warping training views using masked depths, excluding unreliable regions entirely. A photometric loss between the Gaussian rendering at the virtual pose and the warped image then propagates geometric correctness into appearance learning while providing additional viewpoint diversity against overfitting.

Training follows a three-stage curriculum on top of the BinocularGS baseline loss: base 3DGS photometric optimization, activation of geometric regularization (iteration 20k on LLFF/DTU), and finally virtual-view appearance supervision (iteration 25k). Loss weights are λmpc=0.1\lambda_{\text{mpc}}=0.1, Σ\boldsymbol{\Sigma}0, Σ\boldsymbol{\Sigma}1.

Experimental results

Evaluations cover LLFF (forward-facing), DTU (object-centric, weakly textured), and Blender (360° object-centric) at 3/6/9 input views (8 for Blender), with image downsampling factors matching prior baselines. Results are averaged over three seeds.

Dataset / setting Metric Best prior ICO-GS
LLFF, 3-view PSNR 21.44 (BinocularGS) 22.20
LLFF, 6-view PSNR 25.20 (ComapGS) 25.37
DTU, 3-view PSNR 20.71 (BinocularGS) 21.77
DTU, 6-view PSNR 24.51 (CoR-GS) 25.09
DTU, 9-view PSNR 27.18 (CoR-GS) 27.19
Blender, 8-view PSNR 25.42 (DropGaussians) 25.56

On DTU at 3 views—the setting the paper emphasizes—the gain over prior art is +1.06 dB, consistent with the abstract's claim of roughly 1.1 dB improvement. On LLFF, ICO-GS ranks first at 3 and 6 views but second to ComapGS at 9 views (26.45 vs. 26.73 dB), and on Blender it attains the best PSNR while SSIM (0.884 vs. NexusGS's 0.893) and LPIPS (0.100 vs. 0.087) trail slightly—a trade-off the authors attribute to prioritizing geometric accuracy over perceptual metrics. Visual comparisons show sharper depth boundaries and cleaner texture recovery in leaf gaps and weakly-textured surfaces.

Ablations on LLFF and DTU (3-view) confirm each component's contribution relative to the BinocularGS baseline (21.44 / 20.71 dB): removing feature-based multi-view photometric consistency costs −0.38 / −0.46 dB; removing CCDF costs −0.34 / −0.52 dB; removing the virtual-view appearance loss costs −0.41 / −0.57 dB; edge-aware smoothness contributes −0.10 dB on DTU. The full model reaches 22.20 / 21.77 dB, indicating the components are largely complementary rather than redundant.

Limitations and open questions

The paper concedes two material limitations. First, virtual view synthesis assumes view-independent appearance; in regions dominated by specular highlights or reflections, warped appearance provides incorrect supervision. The authors note that sparsity makes these regions difficult for all compared methods, but no mechanism in ICO-GS addresses view-dependence, leaving open how cycle-filtered geometry could be combined with view-dependent appearance modeling. Second, the added regularization increases training time to approximately 1.5× the baseline (~0.3 GB extra memory on 3-view LLFF scenes); whether the curriculum schedule and loss weights transfer without retuning to other capture configurations is not established. Additionally, since the framework builds directly on BinocularGS and inherits its binocular consistency loss, the extent to which gains derive from the new components versus the strong baseline is partially entangled, though the ablation table isolates individual contributions.

Conclusion

ICO-GS frames sparse-view 3DGS degradation as a violation of intrinsic geometry-appearance consistency and addresses it with two coupled mechanisms: occlusion- and illumination-robust multi-view feature-based photometric regularization with edge-preserving depth smoothing, and geometry-guided appearance optimization over virtual views synthesized exclusively from cycle-consistency-filtered depth. Without external monocular depth priors, the method achieves state-of-the-art PSNR across most evaluated settings, with its largest margin (+1.06 dB) precisely on the weakly-textured DTU benchmark at 3 views. The remaining open problems—view-dependent appearance in virtual supervision and computational overhead—are acknowledged by the authors and remain unresolved.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.