Papers
Topics
Authors
Recent
Search
2000 character limit reached

Benchmarking Single-Step Inpainting Methods for Multi-Object 3D Gaussian Splatting Scenes

Published 29 May 2026 in cs.CV | (2605.30987v1)

Abstract: The tasks of object removal and inpainting 3D Gaussian Splatting (3DGS) scenes face challenges such as 3D consistency across camera views. In comparing 2D inpainters and their suitability for the 3D domain, we find that reconstruction-based inpainters outperform generative diffusion models in 3D consistency. Integrating these 2D inpainters into different single-step methods for creating and finetuning 3DGS scenes, our results indicate that initializing the scene from scratch produces higher quality results than finetuning the existing scene. Using a state-of-the-art generative 2D inpainter, we create a straightforward baseline to underline the importance of object removal before inpainting in the 3D setting. Since 360° datasets rarely include real-world ground truths, and challenging occlusion scenarios are equally sparse, we introduce a novel multi-object scene with recorded ground truth data and many views with object occlusions.

Summary

  • The paper benchmarks LaMa, PowerPaint, and Nano Banana within three 3D Gaussian Splatting pipelines, finding that full scene re-initialization outperforms finetuning and that Init-LaMa achieves the strongest overall performance with 19.41 PSNR, 0.7301 SSIM, and 0.3767 LPIPS.
  • Single-step 3D inpainting prioritizes cross-view consistency over individual image quality, allowing reconstruction-based LaMa to outperform sharper generative models such as PowerPaint despite producing less detailed 2D completions.
  • The paper introduces a challenging 360° Living Room scene with ground truth and shows that explicit object removal before hole inpainting is essential for recovering content hidden behind removed objects across multiple views.

Overview

This paper presents a systematic benchmark of single-step inpainting methods for object removal in 3D Gaussian Splatting (3DGS) scenes, conducted by Dröge, Curreli, Saroha, and Cremers of the Technical University of Munich and the Munich Center for Machine Learning (2605.30987). The central question is how effectively 2D inpainting models—ranging from reconstruction-based LaMa to generative diffusion models (PowerPaint, BrushNet) and a state-of-the-art text-to-image model (Nano Banana, i.e., Gemini 2.5 Flash Image)—can be translated into multi-view-consistent 3D completions. The authors introduce three pipeline variants built on Gaussian Grouping and 3DGIC, demonstrate that scene initialization from scratch outperforms finetuning, and contribute a new 360° multi-object "Living Room" scene with recorded ground truth and challenging occlusions.

Pipeline design and method taxonomy

The general pipeline comprises five stages: (1) 3DGS initialization with per-Gaussian semantic identity via Gaussian Grouping, (2) removal of Gaussians belonging to the target object, (3) creation of hole-focused inpainting masks using the 3DGIC method, (4) 2D inpainting of RGB and depth images, and (5) 3D scene completion. The final stage admits two variants: finetuning the existing scene by adding new Gaussians to fill the hole while freezing original ones, or re-initializing the entire scene from the inpainted images—a variant the authors identify as novel among single-step approaches.

Finetuning optimizes only newly added Gaussians with a masked L1\mathcal{L}_1 loss on reference views and an LPIPS loss on non-reference views, plus a masked depth loss and the Gaussian Grouping losses. Notably, the authors deliberately simplify prior work by omitting projection-based cross-view consistency losses, positioning their study as an evaluation of single-step behavior rather than a pursuit of state-of-the-art fidelity.

Methods are denoted as Finetune (FT) or Init with the 2D inpainter (LaMa/LM, PowerPaint/PP, NanoBanana/NB) and, for finetuning, the number of reference views (3 or All). A diagnostic baseline, Init-NanoBanana, inpaints the object directly in all input images without masks and rebuilds the scene—deliberately skipping explicit object removal to test its necessity.

2D inpainter evaluation

Qualitative comparison reveals a fundamental tension: generative models produce higher-fidelity 2D inpaintings but are more fragile operationally. LaMa yields smooth, detail-poor completions that are easier to keep consistent across views; it requires dilated masks significantly larger than the hole but is prompt-free. PowerPaint produces sharper results at the cost of hallucinations and background degradation, partially mitigated by compositing unmasked regions back from the input image—which introduces visible tone discontinuities. Nano Banana inpaints the object rather than the hole when given the raw image, meaning content occluded behind the removed object is never reconstructed. BrushNet performs poorly for object removal—it inserts new foreground objects rather than plain backgrounds, even hallucinating a literal mask because the word appeared in the prompt—and is excluded from 3D evaluation.

In depth inpainting, LaMa excels because the task demands smooth completion matching surrounding geometry, whereas PowerPaint's tone mismatch between input and generated regions is clearly visible in the depth output.

Quantitative results

Averaged over three scenes, Init-LaMa achieves the best overall scores on most metrics:

Method PSNR m-PSNR SSIM LPIPS FID
FT-LM (3) 17.84 18.41 0.688 0.4136 126.12
FT-LM (All) 18.76 19.59 0.7103 0.4016 142.88
FT-PP (All) 15.35 18.01 0.5575 0.5392 291.77
Init-LM 19.41 19.31 0.7301 0.3767 117.88
Init-PP 19.68 19.44 0.7279 0.3814 123.57
Init-NB 17.88 19.80 0.5865 0.4408 131.86

Three findings stand out. First, Init-2DInpainter outperforms both finetuning variants regardless of the 2D inpainter, supporting the claim that full re-initialization is preferable to local hole filling. Second, LaMa-based methods consistently beat PowerPaint-based ones in 3D despite inferior 2D quality, implying that cross-view consistency—not per-image fidelity—is the binding constraint in single-step 3D inpainting. Third, Init-NanoBanana exhibits a sharp dissociation: poor whole-image metrics but competitive masked metrics (best m-FID of 153.82), confirming that its failures stem not from inpainting quality within the mask but from inaccurate content behind the removed object.

Runtime analysis shows finetuning is faster than initialization (since Gaussian Grouping need not be rerun), with Init-NanoBanana fastest at ~1.5 hours—though its external API inference time is excluded, an acknowledged measurement gap. Initialization also yields fewer Gaussians (~860k vs. ~940k).

Qualitative findings and failure modes

Per-scene results expose distinct failure mechanisms. In the bear scene, generative models' view-varying details blend into smooth surfaces under 3D averaging, so all methods converge visually; Init-NanoBanana additionally produces blurry background artifacts from hallucinated occluded content. In the kitchen scene, 2D quality transfers faithfully, and Init-NanoBanana achieves the best result while other methods fail to fully close the hole. In the living room scene—the paper's most demanding test—Init-NanoBanana's lack of 3D awareness is decisive: objects fully occluded by the plush in a given view (a black box, an apple, a pencil sharpener) vanish entirely from the 3D scene, since no input image ever depicts them in that region. This directly demonstrates why object removal followed by hole inpainting, which can exploit cross-view information about occluded content, is necessary.

Supplementary ablations attribute finetuning background blur to 3DGIC's manual foreground depth range, which collapses background features onto a single plane and misplaces background Gaussians closer to the camera. A COLMAP ablation shows that re-running structure-from-motion on inpainted images causes blur (fewer distinctive features after object removal), justifying reuse of original camera parameters—at the cost of residual depth artifacts where the removed object stood.

Limitations and open questions

The authors concede several constraints. Their results do not match 3DGIC's reported quality, attributed to the single-step restriction and omission of cross-view consistent losses—so conclusions here apply specifically to the single-step regime, not iterative refinement pipelines. The evaluation covers only three scenes, one of which is the authors' own contribution, and the benchmark's generality beyond these scenes remains untested. Nano Banana's API runtime cannot be measured, biasing runtime comparisons in its favor. The finding that initialization beats finetuning may interact with the frozen-Gaussian design of the finetuning baseline; whether finetuning with unfrozen background Gaussians would close the gap is left open. Finally, the depth-range trade-off between hole fidelity and background accuracy in finetuning is diagnosed but not resolved.

Conclusion

The paper establishes two clear empirical conclusions for single-step 3DGS inpainting: initializing the scene from scratch on inpainted views outperforms finetuning the existing scene, and reconstruction-based LaMa surpasses stronger generative 2D inpainters once multi-view consistency is required. The Living Room scene with recorded ground truth addresses the scarcity of real-world 360° evaluation data and provides a concrete demonstration that object removal must precede inpainting whenever occluded content matters. The open question the work leaves is whether hybrid schemes—combining generative 2D fidelity with explicit cross-view consistency mechanisms—can retain the sharpness of models like Nano Banana without sacrificing 3D coherence.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.