Papers
Topics
Authors
Recent
Search
2000 character limit reached

3D Gaussian Inpainting Framework

Updated 19 July 2026
  • 3D Gaussian Inpainting is a technique that reconstructs missing scene regions by integrating explicit Gaussian geometry with 2D diffusion and depth completion cues.
  • The framework employs strategies like depth-guided initialization, reference view warping, and bidirectional pipelines to maintain cross-view photorealistic consistency.
  • Empirical studies indicate that optimizing both 2D inpainting and 3D reconstruction leads to enhanced structural accuracy and visual coherence across multiple views.

3D Gaussian inpainting denotes a family of methods that complete missing, occluded, or intentionally removed regions in scenes represented by 3D Gaussian Splatting (3DGS), with the objective that the edited scene remain coherent under novel-view rendering. Within recent work, the term covers several closely related settings: object removal and scene completion, sparse-view completion, 360° multi-object editing, dynamic head editing, and domain-specific reconstruction pipelines that train Gaussian scenes from inpainted supervision rather than from occlusion-corrupted imagery. The central methodological pattern is to couple explicit Gaussian scene optimization with 2D priors—most often diffusion-based inpainting, depth completion, or multiview warping—so that semantic completion in image space is converted into a stable 3D representation (Huang et al., 17 Feb 2025, Liu et al., 2024, Seo et al., 11 Jul 2025, Kim et al., 25 Jun 2026).

1. Problem formulation and representational basis

A 3D Gaussian inpainting framework typically starts from posed multi-view images, camera parameters, and masks that identify the target region to remove or edit. The scene is represented as a set of Gaussians with geometric and appearance parameters, and rendering follows front-to-back alpha compositing. One representative formulation writes the pixel color as

C(p)=∑i=1∣Gp∣ci αi′∏j=1i−1(1−αj′),C(p) = \sum_{i=1}^{|\mathcal{G}_p|} \mathbf{c}_i\, \alpha_i' \prod_{j=1}^{i-1} \left(1 - \alpha_j'\right),

which makes inpainting a problem of introducing or updating Gaussians so that the completed scene remains photometrically plausible from many viewpoints (Dahaghin et al., 9 Sep 2025).

The technical difficulty is not merely to synthesize a plausible image patch. In 3DGS, the missing content must be represented by Gaussians with rendering-relevant properties such as 3D position, opacity, color, and scale or covariance, and the initial 3D positions strongly affect reconstruction quality. This is why several frameworks emphasize geometry-aware initialization, depth completion, or projection-based supervision rather than relying only on RGB completion (Liu et al., 2024).

A second defining constraint is cross-view consistency. Per-view 2D inpainting can produce acceptable single images while still yielding blur, floating structure, depth mismatch, or contradictory textures once the results are fused into a 3D Gaussian field. Recent frameworks therefore treat multi-view agreement, geometry validity, and background preservation as first-class objectives rather than as by-products of reconstruction (Huang et al., 17 Feb 2025, Zhou et al., 24 Jul 2025).

2. Canonical pipeline structures

A recurring pipeline begins with scene initialization, object-level removal, hole-mask construction, 2D image and depth inpainting, and 3D scene completion. In the benchmark of single-step methods, this appears explicitly as a five-stage process: 3D scene initialization, object removal in 3D, inpainting mask creation, 2D image and depth inpainting, and 3D scene completion. That study further distinguishes finetune the existing scene from initialize from scratch, and reports that single-step 2D inpainting followed by scene re-initialization from scratch is generally better than finetuning for 3D consistency (Dröge et al., 29 May 2026).

Other frameworks reorganize the same ingredients around different anchors. Reference-guided systems select one user-chosen or algorithmically favored view and make that view the appearance target for the entire reconstruction. Depth-guided systems first inpaint or complete depth, unproject the result, and then merge or initialize new Gaussians from that geometry. Bidirectional systems alternate between 2D inpainting and 3D Gaussian optimization so that the current 3D scene can guide the next round of image-space completion (Seo et al., 11 Jul 2025, Liu et al., 2024, Kim et al., 25 Jun 2026).

Framework Pipeline form Distinctive mechanism
3DGIC (Huang et al., 17 Feb 2025) Depth-guided mask inference → inpainting-guided 3DGS refinement rendered-depth projection of visible background across views
InFusion (Liu et al., 2024) Remove Gaussians → inpaint reference RGB and complete depth → unproject completed depth and fine-tune diffusion-prior depth completion
RePaintGS (Seo et al., 11 Jul 2025) Initial 3DGS → inpainting confidence evaluation → Inpaint-3DGS reference-guided warping and confidence-weighted view fusion
CoIn (Kim et al., 25 Jun 2026) 2D → 3D → 2D → 3D Reference Adaptive GS with Feature Attention and GS-based Reference Feature Warping
Inpaint360GS (Wang et al., 9 Nov 2025) 3D identity distillation → virtual views → recursive conditional inpainting → depth-guided 3D inpainting object-aware 360° editing

Not all frameworks rely on 2D priors in the same way. GS-RoadPatching instead performs substitutional inpainting directly in 3DGS space by searching for similar existing patches inside the same scene and then applying substitution-and-fusion optimization, while MagicRoad reconstructs a road-centric Gaussian representation from segmentation-guided video inpainting and semantic-aware color enhancement (Chen et al., 24 Sep 2025, Peng et al., 31 Jul 2025).

3. Mechanisms for view, geometry, and temporal consistency

Cross-view consistency is enforced by several non-equivalent mechanisms. In 3DGIC, the key step is depth-guided mask refinement: background pixels visible in one view are projected into another view so that pixels inside an object mask can be removed from the inpainting mask if they are already seen elsewhere as background. This shrinks the unknown region and prevents the method from hallucinating content where the scene is already observed (Huang et al., 17 Feb 2025).

Reference-guided frameworks make consistency explicit in the optimization objective. RePaintGS warps the inpainted reference view into target views, evaluates similarity after geometric alignment, and defines an inpainting weight

winp=σ(α(conf−β)),w_{inp}=\sigma(\alpha(\text{conf}-\beta)),

so that views more consistent with the reference have greater influence during Gaussian optimization. The same framework uses warped-reference pseudo-ground truth only on geometrically consistent pixels and supplements depth with a normal consistency term, thereby coupling appearance transfer with geometry validation (Seo et al., 11 Jul 2025).

A different line of work assigns reliability spatially rather than by whole view. High-fidelity 3D Gaussian Inpainting introduces an automatic Mask Refinement Process and region-wise Uncertainty-guided Optimization. Its uncertainty map is initialized from inpainted depth and converted into a confidence weight map by inversion, so high-confidence blocks exert stronger supervision and less reliable regions are downweighted. This reduces contradictions between sparse inpainted key views while preserving sharper local details (Zhou et al., 24 Jul 2025).

Perspective-aware consistency can also be imposed before Gaussian fitting. PAInpainter builds a perspective graph from LoFTR correspondences, samples anchor and adjacent views from that graph, propagates inpainted anchor content into nearby views via depth-based reprojection, and then selects among multiple diffusion outputs using a dual-feature score

S=ηSrgb+(1−η)Sdepth,S = \eta S_{rgb} + (1-\eta) S_{depth},

with η=0.7\eta = 0.7. The explicit purpose is to reject locally plausible but cross-view inconsistent samples before they are absorbed by the 3D scene (Cheng et al., 13 Oct 2025).

When the target is dynamic rather than static, temporal consistency becomes an additional axis. Edit3DGS applies multi-view batch editing and a lightweight inpainting strategy that isolates latent vectors corresponding to the eyes and mouth and re-injects them later in denoising, addressing the failure mode in which instruction-guided diffusion generates overly similar faces across frames and washes out expression detail (Tran et al., 16 Jun 2026).

4. Major methodological families and extensions

One prominent family is depth-guided Gaussian initialization. InFusion inpaints a reference RGB view, completes its depth with a latent diffusion model, unprojects the completed RGB-D pair into 3D, merges the resulting point cloud into the existing Gaussian scene, and then fine-tunes for only 50–150 iterations. Its central claim is that more accurate depth completion yields better point initialization, which improves both fidelity and efficiency (Liu et al., 2024). RI3D pursues the same principle in sparse-view reconstruction, but separates visible-region repair from missing-region hallucination through two personalized diffusion models and initializes Gaussians from fused depth that combines DUSt3R consistency with monocular relative detail (Paliwal et al., 13 Mar 2025). SplatFill likewise treats depth as the primary geometry prior, using Soft Depth Clustering Loss and Crop-Focused Depth Loss, together with object-aware supervision and Selective Guided Inpainting, to correct only the most inconsistent regions (Dahaghin et al., 9 Sep 2025).

A second family is object-aware or scene-aware reconstruction in challenging camera regimes. Inpaint360GS distills 2D segmentation into stable Gaussian identities with a Key Object Database and Gaussian Set Intersection-over-Union, generates virtual views around the removed object to expose never-before-seen regions, and then performs recursive conditional 2D inpainting followed by depth-guided 3D Gaussian inpainting (Wang et al., 9 Nov 2025). CoIn broadens the task beyond removal by supporting arbitrary-shaped masks and object insertion through a bidirectional 2D → 3D → 2D → 3D pipeline, where diffusion-generated views initialize a coarse 3DGS scene, GS-based warping guides denoising, and a Texture-Enhancing Discriminator restores high-frequency realism (Kim et al., 25 Jun 2026).

A third family replaces RGB-only completion with richer physical or semantic structure. GOR-IS decomposes each Gaussian into intrinsic components including diffuse reflection, Fresnel term, roughness, and object label, performs inpainting in intrinsic space, and uses a lighting-aware mask to suppress residual reflections from the removed object. This formulation is intended for physically consistent object removal in scenes with global lighting effects and non-Lambertian surfaces (Zhao et al., 1 May 2026). VISTA combines visibility-uncertainty-guided 3D Gaussian inpainting with scene conceptual learning, using uncertainty both to weight 3D reconstruction and to learn a textual-inversion scene concept for diffusion-based filling when complementary multiview evidence is insufficient (Cui et al., 23 Apr 2025).

A fourth family is domain-specific direct manipulation in Gaussian space. GS-RoadPatching is designed for driving scenes with repeated structure and avoids 2D generative priors by searching for similar patches in a feature-embedded 3DGS scene and then applying substitution-and-fusion optimization (Chen et al., 24 Sep 2025). MagicRoad is road-surface specific: it removes occluders by segmentation-guided video inpainting, performs semantic-aware color enhancement in HSV space, and then trains a planar-adapted 2D Gaussian surfel representation for clean road reconstruction (Peng et al., 31 Jul 2025).

Closely related extensions treat Gaussian inpainting as a general multiview editing primitive. Instant3dit reframes 3D editing as multiview image inpainting and reconstructs the output back into meshes, NeRFs, or Gaussian Splats in approximately 3 seconds without SDS optimization (Barda et al., 2024). GaussVideoDreamer combines geometry-aware initialization, Inconsistency-Aware Gaussian Splatting, and progressive video inpainting so that rendered views, learned inconsistency masks, and video diffusion form a closed refinement loop (Hao et al., 14 Apr 2025). GauStudio does not propose a direct inpainting algorithm, but its modular initialization, optimization, enhancement, compression, masking, and hybrid foreground–sky modeling make it an enabling framework for completion-oriented Gaussian pipelines (Ye et al., 2024).

5. Benchmarks and empirical evidence

The empirical literature does not converge on one universally dominant recipe, but several trends recur. The benchmark of single-step inpainting methods reports that reconstruction-based LaMa is more 3D-consistent than higher-fidelity generative models such as PowerPaint and Nano Banana, and that initializing the scene from scratch generally yields better reconstruction and inpainting quality than adding Gaussians on top of the old scene (Dröge et al., 29 May 2026). This result is important because it separates per-image visual sharpness from multi-view geometric stability.

Representative reported results illustrate the diversity of current evaluation settings.

Method Dataset or context Reported result
3DGIC + LDM (Huang et al., 17 Feb 2025) SPIn-NeRF FID 36.4, m-FID 96.3, LPIPS 0.26, m-LPIPS 0.028
High-fidelity 3D Gaussian Inpainting (Zhou et al., 24 Jul 2025) reported table LPIPS 0.22, FID 55.17, Time 3 min
SplatFill (Dahaghin et al., 9 Sep 2025) SPIn-NeRF PSNR 20.46, SSIM 0.63, LPIPS 0.25, FID 29.76, Training Time 34 min
PAInpainter (Cheng et al., 13 Oct 2025) SPIn-NeRF / NeRFiller PSNR 26.03 dB / 29.51 dB
Inpaint360GS (Wang et al., 9 Nov 2025) Inpaint360GS dataset PSNR 24.40, SSIM 0.8370, LPIPS 0.1300, masked LPIPS 0.0078, FID 35.93

Several papers tie these numbers to targeted ablations rather than to a single global metric. RePaintGS reports broadly competitive CLIP-based scores against GaussianAvatar-Editor, but its ablation shows that disabling the inpainting module causes edited views to fail to retain the original expression, especially around the eyes and mouth, which worsens animation performance (Tran et al., 16 Jun 2026). RePaintGS also reports that confidence-weighted view fusion outperforms uniform weighting and thresholded filtering, supporting soft reliability weighting rather than naïve aggregation (Seo et al., 11 Jul 2025). PAInpainter’s ablation shows a progression from 27.62 PSNR for a baseline iterative 3DGS system to 29.51 for the full method, with graph sampling, content propagation, and consistency verification all contributing (Cheng et al., 13 Oct 2025). SplatFill reports that its full depth loss improves masked-region performance over both no depth supervision and a GScream-style depth loss, and that Selective Guided Inpainting improves geometry and structural coherence over a single-reference strategy (Dahaghin et al., 9 Sep 2025).

6. Limitations, misconceptions, and research directions

A common misconception is that sharper 2D inpainting necessarily produces better 3D inpainting. The single-step benchmark argues the opposite in many cases: generative diffusion models can look better in 2D, but reconstruction-based inpainters are more stable in 3D because diffusion models tend to hallucinate details that vary across views, whereas smoother outputs are easier to reconcile during Gaussian reconstruction (Dröge et al., 29 May 2026). A related misconception is that multi-view consistency can be recovered simply by perceptual blending of independently inpainted views. RePaintGS explicitly contrasts its reference-guided, confidence-weighted construction with earlier strategies that smooth over inconsistencies but lose fine detail and fail when semantics diverge across views (Seo et al., 11 Jul 2025).

Current frameworks remain constrained by depth quality, view coverage, and scene complexity. RePaintGS can fail around abrupt depth discontinuities, with hallucinations appearing when geometry is highly inconsistent or when there are insufficient high-confidence views for a direction (Seo et al., 11 Jul 2025). High-fidelity 3D Gaussian Inpainting notes difficulty with large and uncontrollable cross-view appearance variation and describes mask refinement as heuristic and scene-dependent (Zhou et al., 24 Jul 2025). Inpaint360GS reports residual shadows, difficulties with irregular complex textures, and occasional floaters in never-before-seen regions (Wang et al., 9 Nov 2025). GOR-IS improves physically consistent removal but still does not explicitly model diffuse global illumination well and struggles with multi-bounce inter-reflections (Zhao et al., 1 May 2026).

The field is also broadening beyond static object removal. CoIn explicitly supports object insertion with flexible mask input, suggesting that future frameworks may treat removal and insertion as instances of a more general bidirectional 2D–3D completion problem (Kim et al., 25 Jun 2026). Dynamic head editing extends masked latent reinjection and multiview editing to temporally varying facial motion (Tran et al., 16 Jun 2026). Video-diffusion-driven methods indicate that inconsistency handling can be formulated as a render–detect–inpaint–update loop rather than as one-shot view aggregation (Hao et al., 14 Apr 2025). GauStudio, although not itself an inpainting method, points toward completion-oriented modularity through customizable initialization, enhancement, masking, and point-cloud completion or upsampling for Gaussian densification (Ye et al., 2024).

Taken together, these developments indicate that a 3D Gaussian inpainting framework is no longer a single design pattern but a research category. The most stable systems combine explicit Gaussian geometry, view-aware reliability estimation, and 2D generative priors, while the most ambitious extensions add temporal modeling, intrinsic decomposition, virtual-view reasoning, or direct 3D patch substitution. This suggests that future progress will depend less on any isolated inpainting backbone than on how effectively 2D semantic completion, 3D geometric constraints, and scene-specific consistency signals are coupled into a single optimization loop.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (18)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to 3D Gaussian Inpainting Framework.