PePESeg3D: Framework for 3D Semantic Segmentation
- PePESeg3D is a two-stage framework for multi-scale 3D semantic segmentation that leverages 3D Gaussian Splatting, utilizing perception priors to improve both geometric reconstruction and feature learning processes, especially enhancing semantic boundary alignment.
- With photometric and depth-constrained geometry reconstruction, the method achieves superior segmentation with an mIoU of 92.3% on SPIn-NeRF and emphasizes the importance of semantic feature learning within a geometrically aligned primitive layout.
- The framework improves semiantic boundary alignment, multi-scale grouping, and novel-view reconstruction using monocular depth, 2D masks, depth-color discontinuities, and view-consistent feature centroids, showing enhanced performance across various benchmarks like LERF-Mask and NVOS.
PePESeg3D is a two-stage framework for multi-scale semantic segmentation with 3D Gaussian Splatting (3DGS) (Choi et al., 23 Sep 2026). It injects perception priors into both Gaussian geometry reconstruction and scale-conditioned feature learning. The method, named PePE Reconstruction and PePE Contrastive Learning, uses monocular relative depth, 2D masks generated by foundation models, depth-color discontinuities, and view-consistent feature centroids to improve semantic boundary alignment, multi-scale grouping, and novel-view reconstruction. Unlike conventional pipelines that first optimize photometric 3DGS geometry and subsequently learn semantic features, PePESeg3D explicitly couples the geometric representation to semantic structure before learning segmentation features.
1. Scope and conceptual foundations
PePESeg3D addresses multi-scale segmentation in scenes represented by explicit 3D Gaussian primitives. A Gaussian has mean , covariance , opacity , and color :
with covariance parameterization
Here, denotes anisotropic scale and denotes rotation. The method attaches a learnable feature vector to each Gaussian and renders color, depth, and semantic features through the same depth-sorted alpha-compositing process used in 3DGS.
Earlier multi-scale 3DGS segmentation methods, including OmniSeg3D-GS and SAGA, generally reconstruct a photometric scene and then learn scale-aware features over that geometry. PePESeg3D identifies two associated limitations. First, photometric reconstruction is not necessarily aware of semantic boundaries: a Gaussian can straddle an object and its background, making it difficult for one primitive to represent both regions with separable semantic features. Second, masks generated by a 2D foundation model such as SAM can be incomplete, inconsistent across viewpoints, and variable in granularity.
The framework therefore uses perception priors twice:
- Geometry alignment: monocular depth and 2D mask boundaries influence Gaussian reconstruction.
- Feature supervision: masks, depth-color cues, and scene-level feature centroids guide scale-conditioned contrastive learning.
The central design principle is that semantic features cannot fully correct a geometrically misaligned primitive layout. A plausible implication is that the quality of the learned segmentation field depends jointly on primitive placement, primitive extent, and feature supervision.
PePESeg3D should be distinguished from PeP, the “Point enhanced Painting” module described in a separate point-cloud paper (Dong et al., 2023). PePESeg3D is an explicitly named 3DGS framework, whereas PeP is a model-agnostic point-cloud feature-enhancement module using point painting and an LM-based point encoder. The two methods address different representations and are not the same architecture.
2. 3DGS representation and complete processing pipeline
For a pixel , the rendered color is obtained by alpha compositing the depth-sorted Gaussians that overlap the pixel:
0
Depth is rendered analogously:
1
where 2 is the view-space depth of Gaussian 3. Semantic features use the same compositing rule.
PePESeg3D operates in two principal stages.
PePE Reconstruction
For each training image, the method obtains an RGB image 4, a set of SAM masks 5, and a monocular relative-depth map 6. Reconstruction proceeds by:
- initializing 3DGS from SfM points;
- optimizing photometric reconstruction together with a monocular-depth loss;
- identifying Gaussians close to visible surfaces;
- refining Gaussians whose projected support crosses mask boundaries;
- continuing standard 3DGS density control and optimization.
The resulting geometry is intended to preserve photometric fidelity while providing better semantic boundary alignment.
PePE Contrastive Learning
After reconstruction, geometry is frozen. A 32-dimensional feature 7 is attached to every Gaussian, together with a learned scale gate 8. The gate produces scale-conditioned features, which are rendered from arbitrary viewpoints and optimized using:
- a mask-based scale-aware contrastive loss;
- a depth-color perception contrastive loss;
- a view-consistent centroid loss;
- feature-norm regularization.
At inference, a query scale produces scale-conditioned feature maps that can support prompt-based segmentation, reference-view label propagation, or automatic clustering with HDBSCAN.
3. PePE Reconstruction and semantic geometry alignment
Gaussian refinement at mask boundaries
The mask-guided refinement stage modifies Gaussians whose projected support crosses a 2D mask boundary. A projected 3D Gaussian has 2D covariance 9. Let 0 be its largest eigenpair. The endpoints of its 1 major-axis ellipse are
2
If this major axis crosses a binary mask boundary, the method finds an intersection point 3 and approximates the fraction of the Gaussian lying inside the mask by
4
The Gaussian is then refined by shrinking it along the major axis and shifting its center:
5
6
Only the mean and scale are replaced; other attributes remain unchanged. Masks are sorted in ascending order of pixel area, so fine masks are processed before coarse masks. Once a Gaussian has been refined, it is excluded from further refinement to prevent repeated shrinking or excessive splitting.
Surface filtering
Not every boundary-crossing Gaussian is refined. Floating or background Gaussians could otherwise be modified incorrectly. For each Gaussian 7, the method computes the depth difference between its projected mean depth 8 and the rendered depth at the corresponding pixel:
9
Only the top 0 of Gaussians with the smallest depth gaps are selected as surface candidates. This preferentially retains Gaussians close to visible surfaces. The appendix generalizes this selection to a refinement ratio 1, with the default 2.
Monocular-depth-constrained optimization
The monocular estimator provides relative rather than metric depth. PePESeg3D therefore aligns rendered depth to the monocular map using optimal scale and shift parameters 3 and 4:
5
The reconstruction objective combines photometric and depth terms:
6
The main implementation uses
7
The optimized variables are the ordinary Gaussian parameters 8, subject to the standard 3DGS densification and pruning schedule. For local depth alignment, the appendix introduces window-specific 9 and 0, allowing local geometric inconsistencies to be captured in addition to global alignment.
For 1 scenes, the default depth weight is 2. For forward-facing scenes, the appendix uses 3 and gradually decays it to 4 later in training, allowing photometric detail refinement to become dominant.
4. Scale-conditioned contrastive feature learning
After reconstruction, each Gaussian receives a learnable feature
5
A two-layer sigmoid-activated gate maps a normalized physical scale 6 to a vector 7. The scale-conditioned Gaussian feature is
8
The scale is derived from the physical extent of a mask. If mask pixels are back-projected into camera coordinates 9, its physical size is
0
Raw scales are normalized using the minimum and maximum scales in each scene. Consequently, the same numerical query scale does not represent the same absolute physical size across different scenes.
Feature rendering and similarity
For query scale 1, rendered features are
2
Pixel-level semantic similarity is cosine similarity:
3
Positive pairs are encouraged toward similarity 4, while negative pairs are pushed toward nonpositive similarity.
Multi-scale mask supervision
For a pixel 5 and query scale 6, the active mask is the finest available mask covering 7 whose physical scale is at least 8:
9
0
Two pixels are treated as belonging to the same entity if their active labels agree or if they occur together in a finer mask:
1
The scale-aware mask loss is
2
The hierarchical condition prevents pixels belonging to a common fine-scale part from becoming dissimilar solely because the queried scale is coarser.
Perception contrastive supervision
Because masks are incomplete, PePESeg3D derives additional dense supervision from RGB and depth. For a query scale 3, it uses a spatial window whose radius increases with scale. Local adaptive thresholds are computed from absolute depth and color gradients:
4
5
Pairs whose depth and color differences are both below their thresholds form 6; pairs whose differences both exceed the thresholds form 7. The perception loss is
8
Depth-color similar pixels are pulled together, while dissimilar pixels are pushed apart. To avoid contradicting reliable mask supervision, pairs for which both pixels lie inside valid mask regions are excluded. Boundary regions are also excluded in the appendix; depth-derived boundaries are preferred over RGB boundaries because RGB gradients are more sensitive to texture and lighting.
View-consistent centroids
Local pairwise losses do not guarantee that the same object receives a consistent identity across views. PePESeg3D therefore maintains centroid sets at two extreme scales:
- upper/coarse scale 9, with centroid set 0;
- lower/fine scale 1, with centroid set 2.
At each scale, candidate feature means from random training views are clustered with HDBSCAN. Candidates whose cosine similarity exceeds 3 are merged, and the resulting centroids are updated using exponential moving averages. Centroids are refreshed every 4 iterations.
For a pixel 5, the nearest centroid is selected by maximum cosine similarity. Only pixels whose similarity exceeds 6 participate in centroid supervision:
7
The centroid loss is
8
The default weights are
9
Centroid losses are activated only after iteration 0.
Feature normalization and total objective
The method regularizes both Gaussian feature norms and rendered feature norms:
1
The final feature-learning objective is
2
with
3
Optional stabilization includes averaging each Gaussian feature with those of its 16 nearest 3D neighbors before rendering, hard-pair oversampling, mask-area reweighting, and periodic centroid refresh.
5. Training, inference, and perception priors
Training procedure
The principal training procedure consists of the following operations:
- Input preparation: camera poses and SfM points are obtained; SAM ViT-H generates masks; Depth-Anything-V2 ViT-B generates relative depth maps.
- 3DGS reconstruction: Gaussian means are initialized from SfM, and photometric and depth-constrained reconstruction is optimized.
- Boundary refinement: beginning at iteration 4, mask-boundary refinement is performed for twice the number of input views.
- Geometry freezing: after reconstruction, geometry and appearance parameters are frozen.
- Feature learning: 32-dimensional Gaussian features and the scale gate are optimized using the contrastive objective.
- Centroid supervision: view-consistent centroids are periodically refreshed and used after iteration 5.
The main settings are:
| Component | Setting | Value |
|---|---|---|
| Reconstruction | Iterations | 30,000 |
| Reconstruction | 6 | 0.2 |
| Reconstruction | 7 | 0.05 |
| Features | Dimension | 32 |
| Feature learning | Iterations | 10,000 |
| Feature learning | Learning rate | 0.0025 |
| Feature learning | Sampled pixels per batch | 1,000 |
| Centroids | Refresh interval | 200 iterations |
| Hardware | GPU | NVIDIA RTX 4090 |
Inference
For a novel view, the frozen Gaussians are projected and alpha-rendered. A query scale 8 is applied through 9, producing scale-conditioned feature maps. The resulting features can be used for:
- prompt-based segmentation by feature similarity;
- label propagation from a reference view;
- automatic decomposition using HDBSCAN.
Increasing the query scale reduces the number of clusters and increases average mask size. Thus, larger values of 0 correspond to progressively coarser grouping, while smaller values support finer object-part segmentation.
Nature of the supervision
SAM ViT-H provides 2D mask proposals, and Depth-Anything-V2 ViT-B provides monocular relative depth. These outputs are treated as noisy and incomplete perception priors rather than complete 3D annotations. Masks constrain primitive boundaries and supply multi-scale pair labels; depth constrains rendered geometry and provides dense boundary-sensitive cues when masks are unavailable.
The appendix reports robustness to replacing SAM with SAM2, which generates 25–30% fewer masks, and to replacing Depth-Anything-V2 with Depth-Anything-V3 or MiDaS.
6. Evaluation, ablations, and computational characteristics
PePESeg3D is evaluated on SPIn-NeRF, LERF-Mask, LERF-Mask-Fine, and NVOS. Segmentation metrics include mIoU, pixel accuracy, and mean Boundary IoU; reconstruction metrics include PSNR, SSIM, and LPIPS.
Benchmark results
On SPIn-NeRF, PePESeg3D achieves the following results:
| Method | mIoU | mAcc | PSNR | SSIM | LPIPS |
|---|---|---|---|---|---|
| SAGA | 91.9 | 98.9 | 26.09 | 0.850 | 0.153 |
| PePESeg3D | 92.3 | 98.9 | 26.81 | 0.857 | 0.142 |
The paper reports a 1-percentage-point mIoU improvement over UnifiedLift, the strongest single-scale baseline in that comparison.
On LERF-Mask:
| Method | mIoU | mBIoU |
|---|---|---|
| SAGA | 78.4 | 74.0 |
| OpenSplat3D | 83.5 | 78.4 |
| PePESeg3D | 80.5 | 76.5 |
Among multi-scale methods, PePESeg3D exceeds SAGA by 2 mIoU points and 3 mBIoU points.
On LERF-Mask-Fine:
| Method | mIoU | mBIoU |
|---|---|---|
| OmniSeg3D-GS | 39.8 | 37.3 |
| SAGA | 69.6 | 67.8 |
| PePESeg3D | 70.8 | 68.3 |
On NVOS segmentation:
| Method | mIoU | mAcc |
|---|---|---|
| SAGA | 91.5 | 98.3 |
| PePESeg3D | 92.2 | 98.5 |
On NVOS reconstruction:
| Method | PSNR | SSIM | LPIPS |
|---|---|---|---|
| SAGA | 25.86 | 0.837 | 0.161 |
| PePESeg3D | 27.08 | 0.850 | 0.145 |
The paper reports a 4-dB PSNR improvement over SAGA on NVOS. Qualitatively, depth supervision reduces floating artifacts, including noisy Gaussians between the T-Rex skeleton and background wall.
Ablation results
On NVOS reconstruction:
| Variant | PSNR | SSIM | LPIPS |
|---|---|---|---|
| Baseline | 25.86 | 0.837 | 0.161 |
| + Gaussian refinement | 26.32 | 0.843 | 0.154 |
| + monocular depth, full reconstruction | 27.08 | 0.850 | 0.145 |
Gaussian refinement provides a 5-dB PSNR gain, while monocular depth adds another 6 dB.
On LERF-Mask segmentation:
| Variant | mIoU | mBIoU |
|---|---|---|
| Baseline SAGA-style system | 78.4 | 74.0 |
| + PePE Reconstruction | 79.2 | 75.3 |
| + perception loss | 80.2 | 76.0 |
| + consistency/centroid loss, full model | 80.5 | 76.5 |
The reported incremental improvements are 7 mIoU from PePE Reconstruction, 8 mIoU from perception loss, and 9 mIoU from view consistency.
On LERF-Mask-Fine, query scale controls the granularity of grouping:
| Query scale | Fine-part mIoU | Clusters | Mean mask size |
|---|---|---|---|
| 0.01 | 68.4 | 83 | 2,336 px |
| 0.3 | 65.8 | 79 | 2,337 px |
| 0.5 | 48.6 | 61 | 3,439 px |
| 0.7 | 27.1 | 49 | 3,910 px |
| 1.0 | 19.2 | 43 | 5,058 px |
This behavior supports the interpretation that increasing 00 produces coarser groups.
Robustness
On LERF-Mask-Fine, default PePESeg3D obtains 01 mIoU/mBIoU. With one-quarter of the views, performance decreases to 02; with SAM2, it becomes 03. The corresponding SAGA results are 04, 05, and 06. Replacing the depth backbone with Depth-Anything-V3 or MiDaS changes performance by at most roughly 07 mIoU and 08 mBIoU.
Computational cost
On SPIn-NeRF using an RTX 4090:
| Method | VRAM | Reconstruction | Segmentation | Inference |
|---|---|---|---|---|
| SAGA | 12.1 GiB | 7.7 min | 15.4 min | 6.8 ms |
| PePESeg3D | 13.1 GiB | 11.2 min | 20.9 min | 6.6 ms |
Relative to SAGA, PePESeg3D uses approximately 09 GiB more VRAM, requires 10 additional reconstruction minutes and 11 additional segmentation minutes, and has comparable inference latency.
The reported PePESeg3D time breakdown is:
- reconstruction shared cost: 9.71 minutes;
- Gaussian refinement: 0.27 minutes;
- depth supervision: 1.27 minutes;
- segmentation shared cost: 15.80 minutes;
- scale-aware loss: 0.62 minutes;
- perception loss: 4.04 minutes;
- centroid loss: 0.34 minutes.
The perception loss is the dominant additional segmentation cost.
7. Limitations and research significance
PePESeg3D has three explicitly identified limitations.
Cross-scale coupling: the scale gate acts on a shared feature space, so channels activated at one scale can also be activated at another. Nearby query scales may therefore produce similar representations, complicating separation of entities with nearly identical physical sizes.
Incomplete segmentation of complicated structures: perception loss cannot fully compensate when SAM fails to recognize a complicated object as a coherent entity. Intricate targets may remain fragmented or incorrectly grouped. Global connectivity reasoning is suggested as a possible remedy.
Scene-dependent scale normalization: scales are normalized independently within each scene. Consequently, the same numerical query scale does not denote the same absolute physical size across scenes. A reference object or scale sweep may be required in a new scene.
The broader significance of the method lies in its bidirectional coupling of geometry and semantics. Monocular depth improves spatial ordering and suppresses floating Gaussians; mask boundaries reduce primitive straddling; depth-color cues fill gaps in 2D mask supervision; centroid constraints improve cross-view identity consistency; and the continuous scale gate supports a hierarchy from fine object parts to coarse whole-object groupings.
PePESeg3D therefore treats multi-scale 3D segmentation as both a representation problem and a supervision problem. Its principal contribution is not merely the addition of contrastive losses, but the construction of a semantically better 3D Gaussian substrate before learning scale-conditioned features. The reported results indicate that geometry refinement, perception-derived contrastive supervision, and view-consistent centroid learning each contribute to segmentation quality, while the reconstruction stage also improves novel-view rendering.