Papers
Topics
Authors
Recent
Search
2000 character limit reached

PePESeg3D: Framework for 3D Semantic Segmentation

Updated 29 September 2026
  • PePESeg3D is a two-stage framework for multi-scale 3D semantic segmentation that leverages 3D Gaussian Splatting, utilizing perception priors to improve both geometric reconstruction and feature learning processes, especially enhancing semantic boundary alignment.
  • With photometric and depth-constrained geometry reconstruction, the method achieves superior segmentation with an mIoU of 92.3% on SPIn-NeRF and emphasizes the importance of semantic feature learning within a geometrically aligned primitive layout.
  • The framework improves semiantic boundary alignment, multi-scale grouping, and novel-view reconstruction using monocular depth, 2D masks, depth-color discontinuities, and view-consistent feature centroids, showing enhanced performance across various benchmarks like LERF-Mask and NVOS.

PePESeg3D is a two-stage framework for multi-scale semantic segmentation with 3D Gaussian Splatting (3DGS) (Choi et al., 23 Sep 2026). It injects perception priors into both Gaussian geometry reconstruction and scale-conditioned feature learning. The method, named PePE Reconstruction and PePE Contrastive Learning, uses monocular relative depth, 2D masks generated by foundation models, depth-color discontinuities, and view-consistent feature centroids to improve semantic boundary alignment, multi-scale grouping, and novel-view reconstruction. Unlike conventional pipelines that first optimize photometric 3DGS geometry and subsequently learn semantic features, PePESeg3D explicitly couples the geometric representation to semantic structure before learning segmentation features.

1. Scope and conceptual foundations

PePESeg3D addresses multi-scale segmentation in scenes represented by explicit 3D Gaussian primitives. A Gaussian has mean μ\boldsymbol{\mu}, covariance Σ\boldsymbol{\Sigma}, opacity σ\boldsymbol{\sigma}, and color c\mathbf c:

g(x)=exp⁡{−12(x−μ)⊤Σ−1(x−μ)},g(\mathbf{x})= \exp\left\{ -\frac12(\mathbf{x}-\boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x}-\boldsymbol{\mu}) \right\},

with covariance parameterization

Σ=RSS⊤R⊤.\boldsymbol{\Sigma}=\mathbf R\mathbf S\mathbf S^\top\mathbf R^\top.

Here, S\mathbf S denotes anisotropic scale and R\mathbf R denotes rotation. The method attaches a learnable feature vector to each Gaussian and renders color, depth, and semantic features through the same depth-sorted alpha-compositing process used in 3DGS.

Earlier multi-scale 3DGS segmentation methods, including OmniSeg3D-GS and SAGA, generally reconstruct a photometric scene and then learn scale-aware features over that geometry. PePESeg3D identifies two associated limitations. First, photometric reconstruction is not necessarily aware of semantic boundaries: a Gaussian can straddle an object and its background, making it difficult for one primitive to represent both regions with separable semantic features. Second, masks generated by a 2D foundation model such as SAM can be incomplete, inconsistent across viewpoints, and variable in granularity.

The framework therefore uses perception priors twice:

  1. Geometry alignment: monocular depth and 2D mask boundaries influence Gaussian reconstruction.
  2. Feature supervision: masks, depth-color cues, and scene-level feature centroids guide scale-conditioned contrastive learning.

The central design principle is that semantic features cannot fully correct a geometrically misaligned primitive layout. A plausible implication is that the quality of the learned segmentation field depends jointly on primitive placement, primitive extent, and feature supervision.

PePESeg3D should be distinguished from PeP, the “Point enhanced Painting” module described in a separate point-cloud paper (Dong et al., 2023). PePESeg3D is an explicitly named 3DGS framework, whereas PeP is a model-agnostic point-cloud feature-enhancement module using point painting and an LM-based point encoder. The two methods address different representations and are not the same architecture.

2. 3DGS representation and complete processing pipeline

For a pixel pp, the rendered color is obtained by alpha compositing the NpN_p depth-sorted Gaussians that overlap the pixel:

Σ\boldsymbol{\Sigma}0

Depth is rendered analogously:

Σ\boldsymbol{\Sigma}1

where Σ\boldsymbol{\Sigma}2 is the view-space depth of Gaussian Σ\boldsymbol{\Sigma}3. Semantic features use the same compositing rule.

PePESeg3D operates in two principal stages.

PePE Reconstruction

For each training image, the method obtains an RGB image Σ\boldsymbol{\Sigma}4, a set of SAM masks Σ\boldsymbol{\Sigma}5, and a monocular relative-depth map Σ\boldsymbol{\Sigma}6. Reconstruction proceeds by:

  1. initializing 3DGS from SfM points;
  2. optimizing photometric reconstruction together with a monocular-depth loss;
  3. identifying Gaussians close to visible surfaces;
  4. refining Gaussians whose projected support crosses mask boundaries;
  5. continuing standard 3DGS density control and optimization.

The resulting geometry is intended to preserve photometric fidelity while providing better semantic boundary alignment.

PePE Contrastive Learning

After reconstruction, geometry is frozen. A 32-dimensional feature Σ\boldsymbol{\Sigma}7 is attached to every Gaussian, together with a learned scale gate Σ\boldsymbol{\Sigma}8. The gate produces scale-conditioned features, which are rendered from arbitrary viewpoints and optimized using:

  • a mask-based scale-aware contrastive loss;
  • a depth-color perception contrastive loss;
  • a view-consistent centroid loss;
  • feature-norm regularization.

At inference, a query scale produces scale-conditioned feature maps that can support prompt-based segmentation, reference-view label propagation, or automatic clustering with HDBSCAN.

3. PePE Reconstruction and semantic geometry alignment

Gaussian refinement at mask boundaries

The mask-guided refinement stage modifies Gaussians whose projected support crosses a 2D mask boundary. A projected 3D Gaussian has 2D covariance Σ\boldsymbol{\Sigma}9. Let σ\boldsymbol{\sigma}0 be its largest eigenpair. The endpoints of its σ\boldsymbol{\sigma}1 major-axis ellipse are

σ\boldsymbol{\sigma}2

If this major axis crosses a binary mask boundary, the method finds an intersection point σ\boldsymbol{\sigma}3 and approximates the fraction of the Gaussian lying inside the mask by

σ\boldsymbol{\sigma}4

The Gaussian is then refined by shrinking it along the major axis and shifting its center:

σ\boldsymbol{\sigma}5

σ\boldsymbol{\sigma}6

Only the mean and scale are replaced; other attributes remain unchanged. Masks are sorted in ascending order of pixel area, so fine masks are processed before coarse masks. Once a Gaussian has been refined, it is excluded from further refinement to prevent repeated shrinking or excessive splitting.

Surface filtering

Not every boundary-crossing Gaussian is refined. Floating or background Gaussians could otherwise be modified incorrectly. For each Gaussian σ\boldsymbol{\sigma}7, the method computes the depth difference between its projected mean depth σ\boldsymbol{\sigma}8 and the rendered depth at the corresponding pixel:

σ\boldsymbol{\sigma}9

Only the top c\mathbf c0 of Gaussians with the smallest depth gaps are selected as surface candidates. This preferentially retains Gaussians close to visible surfaces. The appendix generalizes this selection to a refinement ratio c\mathbf c1, with the default c\mathbf c2.

Monocular-depth-constrained optimization

The monocular estimator provides relative rather than metric depth. PePESeg3D therefore aligns rendered depth to the monocular map using optimal scale and shift parameters c\mathbf c3 and c\mathbf c4:

c\mathbf c5

The reconstruction objective combines photometric and depth terms:

c\mathbf c6

The main implementation uses

c\mathbf c7

The optimized variables are the ordinary Gaussian parameters c\mathbf c8, subject to the standard 3DGS densification and pruning schedule. For local depth alignment, the appendix introduces window-specific c\mathbf c9 and g(x)=exp⁡{−12(x−μ)⊤Σ−1(x−μ)},g(\mathbf{x})= \exp\left\{ -\frac12(\mathbf{x}-\boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x}-\boldsymbol{\mu}) \right\},0, allowing local geometric inconsistencies to be captured in addition to global alignment.

For g(x)=exp⁡{−12(x−μ)⊤Σ−1(x−μ)},g(\mathbf{x})= \exp\left\{ -\frac12(\mathbf{x}-\boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x}-\boldsymbol{\mu}) \right\},1 scenes, the default depth weight is g(x)=exp⁡{−12(x−μ)⊤Σ−1(x−μ)},g(\mathbf{x})= \exp\left\{ -\frac12(\mathbf{x}-\boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x}-\boldsymbol{\mu}) \right\},2. For forward-facing scenes, the appendix uses g(x)=exp⁡{−12(x−μ)⊤Σ−1(x−μ)},g(\mathbf{x})= \exp\left\{ -\frac12(\mathbf{x}-\boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x}-\boldsymbol{\mu}) \right\},3 and gradually decays it to g(x)=exp⁡{−12(x−μ)⊤Σ−1(x−μ)},g(\mathbf{x})= \exp\left\{ -\frac12(\mathbf{x}-\boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x}-\boldsymbol{\mu}) \right\},4 later in training, allowing photometric detail refinement to become dominant.

4. Scale-conditioned contrastive feature learning

After reconstruction, each Gaussian receives a learnable feature

g(x)=exp⁡{−12(x−μ)⊤Σ−1(x−μ)},g(\mathbf{x})= \exp\left\{ -\frac12(\mathbf{x}-\boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x}-\boldsymbol{\mu}) \right\},5

A two-layer sigmoid-activated gate maps a normalized physical scale g(x)=exp⁡{−12(x−μ)⊤Σ−1(x−μ)},g(\mathbf{x})= \exp\left\{ -\frac12(\mathbf{x}-\boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x}-\boldsymbol{\mu}) \right\},6 to a vector g(x)=exp⁡{−12(x−μ)⊤Σ−1(x−μ)},g(\mathbf{x})= \exp\left\{ -\frac12(\mathbf{x}-\boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x}-\boldsymbol{\mu}) \right\},7. The scale-conditioned Gaussian feature is

g(x)=exp⁡{−12(x−μ)⊤Σ−1(x−μ)},g(\mathbf{x})= \exp\left\{ -\frac12(\mathbf{x}-\boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x}-\boldsymbol{\mu}) \right\},8

The scale is derived from the physical extent of a mask. If mask pixels are back-projected into camera coordinates g(x)=exp⁡{−12(x−μ)⊤Σ−1(x−μ)},g(\mathbf{x})= \exp\left\{ -\frac12(\mathbf{x}-\boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x}-\boldsymbol{\mu}) \right\},9, its physical size is

Σ=RSS⊤R⊤.\boldsymbol{\Sigma}=\mathbf R\mathbf S\mathbf S^\top\mathbf R^\top.0

Raw scales are normalized using the minimum and maximum scales in each scene. Consequently, the same numerical query scale does not represent the same absolute physical size across different scenes.

Feature rendering and similarity

For query scale Σ=RSS⊤R⊤.\boldsymbol{\Sigma}=\mathbf R\mathbf S\mathbf S^\top\mathbf R^\top.1, rendered features are

Σ=RSS⊤R⊤.\boldsymbol{\Sigma}=\mathbf R\mathbf S\mathbf S^\top\mathbf R^\top.2

Pixel-level semantic similarity is cosine similarity:

Σ=RSS⊤R⊤.\boldsymbol{\Sigma}=\mathbf R\mathbf S\mathbf S^\top\mathbf R^\top.3

Positive pairs are encouraged toward similarity Σ=RSS⊤R⊤.\boldsymbol{\Sigma}=\mathbf R\mathbf S\mathbf S^\top\mathbf R^\top.4, while negative pairs are pushed toward nonpositive similarity.

Multi-scale mask supervision

For a pixel Σ=RSS⊤R⊤.\boldsymbol{\Sigma}=\mathbf R\mathbf S\mathbf S^\top\mathbf R^\top.5 and query scale Σ=RSS⊤R⊤.\boldsymbol{\Sigma}=\mathbf R\mathbf S\mathbf S^\top\mathbf R^\top.6, the active mask is the finest available mask covering Σ=RSS⊤R⊤.\boldsymbol{\Sigma}=\mathbf R\mathbf S\mathbf S^\top\mathbf R^\top.7 whose physical scale is at least Σ=RSS⊤R⊤.\boldsymbol{\Sigma}=\mathbf R\mathbf S\mathbf S^\top\mathbf R^\top.8:

Σ=RSS⊤R⊤.\boldsymbol{\Sigma}=\mathbf R\mathbf S\mathbf S^\top\mathbf R^\top.9

S\mathbf S0

Two pixels are treated as belonging to the same entity if their active labels agree or if they occur together in a finer mask:

S\mathbf S1

The scale-aware mask loss is

S\mathbf S2

The hierarchical condition prevents pixels belonging to a common fine-scale part from becoming dissimilar solely because the queried scale is coarser.

Perception contrastive supervision

Because masks are incomplete, PePESeg3D derives additional dense supervision from RGB and depth. For a query scale S\mathbf S3, it uses a spatial window whose radius increases with scale. Local adaptive thresholds are computed from absolute depth and color gradients:

S\mathbf S4

S\mathbf S5

Pairs whose depth and color differences are both below their thresholds form S\mathbf S6; pairs whose differences both exceed the thresholds form S\mathbf S7. The perception loss is

S\mathbf S8

Depth-color similar pixels are pulled together, while dissimilar pixels are pushed apart. To avoid contradicting reliable mask supervision, pairs for which both pixels lie inside valid mask regions are excluded. Boundary regions are also excluded in the appendix; depth-derived boundaries are preferred over RGB boundaries because RGB gradients are more sensitive to texture and lighting.

View-consistent centroids

Local pairwise losses do not guarantee that the same object receives a consistent identity across views. PePESeg3D therefore maintains centroid sets at two extreme scales:

  • upper/coarse scale S\mathbf S9, with centroid set R\mathbf R0;
  • lower/fine scale R\mathbf R1, with centroid set R\mathbf R2.

At each scale, candidate feature means from random training views are clustered with HDBSCAN. Candidates whose cosine similarity exceeds R\mathbf R3 are merged, and the resulting centroids are updated using exponential moving averages. Centroids are refreshed every R\mathbf R4 iterations.

For a pixel R\mathbf R5, the nearest centroid is selected by maximum cosine similarity. Only pixels whose similarity exceeds R\mathbf R6 participate in centroid supervision:

R\mathbf R7

The centroid loss is

R\mathbf R8

The default weights are

R\mathbf R9

Centroid losses are activated only after iteration pp0.

Feature normalization and total objective

The method regularizes both Gaussian feature norms and rendered feature norms:

pp1

The final feature-learning objective is

pp2

with

pp3

Optional stabilization includes averaging each Gaussian feature with those of its 16 nearest 3D neighbors before rendering, hard-pair oversampling, mask-area reweighting, and periodic centroid refresh.

5. Training, inference, and perception priors

Training procedure

The principal training procedure consists of the following operations:

  1. Input preparation: camera poses and SfM points are obtained; SAM ViT-H generates masks; Depth-Anything-V2 ViT-B generates relative depth maps.
  2. 3DGS reconstruction: Gaussian means are initialized from SfM, and photometric and depth-constrained reconstruction is optimized.
  3. Boundary refinement: beginning at iteration pp4, mask-boundary refinement is performed for twice the number of input views.
  4. Geometry freezing: after reconstruction, geometry and appearance parameters are frozen.
  5. Feature learning: 32-dimensional Gaussian features and the scale gate are optimized using the contrastive objective.
  6. Centroid supervision: view-consistent centroids are periodically refreshed and used after iteration pp5.

The main settings are:

Component Setting Value
Reconstruction Iterations 30,000
Reconstruction pp6 0.2
Reconstruction pp7 0.05
Features Dimension 32
Feature learning Iterations 10,000
Feature learning Learning rate 0.0025
Feature learning Sampled pixels per batch 1,000
Centroids Refresh interval 200 iterations
Hardware GPU NVIDIA RTX 4090

Inference

For a novel view, the frozen Gaussians are projected and alpha-rendered. A query scale pp8 is applied through pp9, producing scale-conditioned feature maps. The resulting features can be used for:

  • prompt-based segmentation by feature similarity;
  • label propagation from a reference view;
  • automatic decomposition using HDBSCAN.

Increasing the query scale reduces the number of clusters and increases average mask size. Thus, larger values of NpN_p0 correspond to progressively coarser grouping, while smaller values support finer object-part segmentation.

Nature of the supervision

SAM ViT-H provides 2D mask proposals, and Depth-Anything-V2 ViT-B provides monocular relative depth. These outputs are treated as noisy and incomplete perception priors rather than complete 3D annotations. Masks constrain primitive boundaries and supply multi-scale pair labels; depth constrains rendered geometry and provides dense boundary-sensitive cues when masks are unavailable.

The appendix reports robustness to replacing SAM with SAM2, which generates 25–30% fewer masks, and to replacing Depth-Anything-V2 with Depth-Anything-V3 or MiDaS.

6. Evaluation, ablations, and computational characteristics

PePESeg3D is evaluated on SPIn-NeRF, LERF-Mask, LERF-Mask-Fine, and NVOS. Segmentation metrics include mIoU, pixel accuracy, and mean Boundary IoU; reconstruction metrics include PSNR, SSIM, and LPIPS.

Benchmark results

On SPIn-NeRF, PePESeg3D achieves the following results:

Method mIoU mAcc PSNR SSIM LPIPS
SAGA 91.9 98.9 26.09 0.850 0.153
PePESeg3D 92.3 98.9 26.81 0.857 0.142

The paper reports a NpN_p1-percentage-point mIoU improvement over UnifiedLift, the strongest single-scale baseline in that comparison.

On LERF-Mask:

Method mIoU mBIoU
SAGA 78.4 74.0
OpenSplat3D 83.5 78.4
PePESeg3D 80.5 76.5

Among multi-scale methods, PePESeg3D exceeds SAGA by NpN_p2 mIoU points and NpN_p3 mBIoU points.

On LERF-Mask-Fine:

Method mIoU mBIoU
OmniSeg3D-GS 39.8 37.3
SAGA 69.6 67.8
PePESeg3D 70.8 68.3

On NVOS segmentation:

Method mIoU mAcc
SAGA 91.5 98.3
PePESeg3D 92.2 98.5

On NVOS reconstruction:

Method PSNR SSIM LPIPS
SAGA 25.86 0.837 0.161
PePESeg3D 27.08 0.850 0.145

The paper reports a NpN_p4-dB PSNR improvement over SAGA on NVOS. Qualitatively, depth supervision reduces floating artifacts, including noisy Gaussians between the T-Rex skeleton and background wall.

Ablation results

On NVOS reconstruction:

Variant PSNR SSIM LPIPS
Baseline 25.86 0.837 0.161
+ Gaussian refinement 26.32 0.843 0.154
+ monocular depth, full reconstruction 27.08 0.850 0.145

Gaussian refinement provides a NpN_p5-dB PSNR gain, while monocular depth adds another NpN_p6 dB.

On LERF-Mask segmentation:

Variant mIoU mBIoU
Baseline SAGA-style system 78.4 74.0
+ PePE Reconstruction 79.2 75.3
+ perception loss 80.2 76.0
+ consistency/centroid loss, full model 80.5 76.5

The reported incremental improvements are NpN_p7 mIoU from PePE Reconstruction, NpN_p8 mIoU from perception loss, and NpN_p9 mIoU from view consistency.

On LERF-Mask-Fine, query scale controls the granularity of grouping:

Query scale Fine-part mIoU Clusters Mean mask size
0.01 68.4 83 2,336 px
0.3 65.8 79 2,337 px
0.5 48.6 61 3,439 px
0.7 27.1 49 3,910 px
1.0 19.2 43 5,058 px

This behavior supports the interpretation that increasing Σ\boldsymbol{\Sigma}00 produces coarser groups.

Robustness

On LERF-Mask-Fine, default PePESeg3D obtains Σ\boldsymbol{\Sigma}01 mIoU/mBIoU. With one-quarter of the views, performance decreases to Σ\boldsymbol{\Sigma}02; with SAM2, it becomes Σ\boldsymbol{\Sigma}03. The corresponding SAGA results are Σ\boldsymbol{\Sigma}04, Σ\boldsymbol{\Sigma}05, and Σ\boldsymbol{\Sigma}06. Replacing the depth backbone with Depth-Anything-V3 or MiDaS changes performance by at most roughly Σ\boldsymbol{\Sigma}07 mIoU and Σ\boldsymbol{\Sigma}08 mBIoU.

Computational cost

On SPIn-NeRF using an RTX 4090:

Method VRAM Reconstruction Segmentation Inference
SAGA 12.1 GiB 7.7 min 15.4 min 6.8 ms
PePESeg3D 13.1 GiB 11.2 min 20.9 min 6.6 ms

Relative to SAGA, PePESeg3D uses approximately Σ\boldsymbol{\Sigma}09 GiB more VRAM, requires Σ\boldsymbol{\Sigma}10 additional reconstruction minutes and Σ\boldsymbol{\Sigma}11 additional segmentation minutes, and has comparable inference latency.

The reported PePESeg3D time breakdown is:

  • reconstruction shared cost: 9.71 minutes;
  • Gaussian refinement: 0.27 minutes;
  • depth supervision: 1.27 minutes;
  • segmentation shared cost: 15.80 minutes;
  • scale-aware loss: 0.62 minutes;
  • perception loss: 4.04 minutes;
  • centroid loss: 0.34 minutes.

The perception loss is the dominant additional segmentation cost.

7. Limitations and research significance

PePESeg3D has three explicitly identified limitations.

Cross-scale coupling: the scale gate acts on a shared feature space, so channels activated at one scale can also be activated at another. Nearby query scales may therefore produce similar representations, complicating separation of entities with nearly identical physical sizes.

Incomplete segmentation of complicated structures: perception loss cannot fully compensate when SAM fails to recognize a complicated object as a coherent entity. Intricate targets may remain fragmented or incorrectly grouped. Global connectivity reasoning is suggested as a possible remedy.

Scene-dependent scale normalization: scales are normalized independently within each scene. Consequently, the same numerical query scale does not denote the same absolute physical size across scenes. A reference object or scale sweep may be required in a new scene.

The broader significance of the method lies in its bidirectional coupling of geometry and semantics. Monocular depth improves spatial ordering and suppresses floating Gaussians; mask boundaries reduce primitive straddling; depth-color cues fill gaps in 2D mask supervision; centroid constraints improve cross-view identity consistency; and the continuous scale gate supports a hierarchy from fine object parts to coarse whole-object groupings.

PePESeg3D therefore treats multi-scale 3D segmentation as both a representation problem and a supervision problem. Its principal contribution is not merely the addition of contrastive losses, but the construction of a semantically better 3D Gaussian substrate before learning scale-conditioned features. The reported results indicate that geometry refinement, perception-derived contrastive supervision, and view-consistent centroid learning each contribute to segmentation quality, while the reconstruction stage also improves novel-view rendering.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to PePESeg3D.