---
title: 'PreSem-Surf: RGB-D NeRF/SDF Reconstruction'
url: https://www.emergentmind.com/topics/presem-surf
type: topic
---

# PreSem-Surf: RGB-D NeRF/SDF Reconstruction

Searching arXiv for PreSem-Surf and closely related papers mentioned in the provided data.
First, locate the primary PreSem-Surf paper.
PreSem-Surf is an RGB-D surface reconstruction method built on NeRF/SDF-style neural implicit representations and designed to reconstruct high-quality meshes of indoor scenes from RGB-D sequences while simultaneously leveraging semantic information and accelerating training. It integrates RGB, depth, and semantic information to improve reconstruction performance, and its defining components are a voxel pre-rendering mechanism based on a Sampling-Guided MLP (SG-MLP) combined with a Preconditioning MLP (PR-MLP), together with a progressive semantic modeling strategy denoted PFPSMS. In the reported experiments on seven synthetic scenes with six evaluation metrics, it achieves the best performance in C-L1, F-score, and IoU, while maintaining competitive results in NC, Accuracy, and Completeness [2508.13228].

## 1. Concept and research setting

PreSem-Surf is situated within the NeRF ecosystem, where a scene is modeled as a continuous function over 3D position and viewing direction and rendered by volumetric integration. In the formulation summarized for PreSem-Surf, standard NeRF is characterized as effective for novel view synthesis but slow to train due to dense ray sampling and large MLPs, sensitive to noisy or incomplete inputs, not inherently semantic, and often inefficient for extracting clean surfaces. The RGB-D setting supplies both color and depth, but the problem remains difficult because depth maps are noisy and incomplete, RGB images contain view-dependent effects and occlusions, and dense sampling imposes heavy memory and compute demands [2508.13228].

The method is explicitly motivated by two goals. The first is sampling and rendering efficiency with high geometric fidelity, implemented through a voxel-based, sampling-guided pre-rendering mechanism that quickly captures global scene structure and then focuses computation on important regions. The second is progressive semantic modeling, implemented as a coarse-to-fine integration of semantics using pseudo-color semantic images and multi-stage rendering [2508.13228].

A plausible implication is that PreSem-Surf treats semantics not as an auxiliary annotation channel appended late in optimization, but as part of the reconstruction schedule itself. This distinguishes it conceptually from pipelines in which geometry is first reconstructed and semantics are attached only afterward.

## 2. Architectural organization

The overall pipeline contains five components. First, a simplified PR-MLP performs coarse volumetric density or SDF-based distance estimation across voxels using uniform sampling, building a rough scaffold of the scene. Second, SG-MLP refines sampling hierarchically, allocating more samples to regions with higher estimated density and fewer to empty or noisy areas. Third, two MLPs are applied to the retained samples: an RGB MLP that predicts color for volume rendering and an SDF MLP that predicts signed distance for surface-based rendering and mesh extraction. Fourth, PFPSMS introduces progressive semantic modeling using pseudo-color semantic images derived from DFormer segmentation and NYU40 color mapping. Fifth, the learned SDF field is converted to an explicit mesh through zero-level-set extraction [2508.13228].

The representation is SDF-based and follows a NEUS-style rendering formulation. The scene is represented by a continuous signed distance function $\text{SDF}(\mathbf{x})$, and density is derived from SDF values by
$$
\sigma_j = \varphi(\text{sdf}_j \cdot \text{inv}_s),
$$
where $\text{sdf}_j$ is the SDF at voxel $j$, $\text{inv}_s$ is a scaling factor, and $\varphi$ is the sigmoid activation. Per-ray weights are then computed as
$$
\omega_k = \exp\left(-\sum_{j=1}^{k-1} \sigma_j \Delta z_j \right) \cdot \left(1 - \exp(-\sigma_k \Delta z_k)\right),
$$
and the rendered color is
$$
C(\mathbf{r}) = \sum_k \omega_k \, \mathbf{c}(\mathbf{x}_k).
$$
Rendered depth is produced by weighted averaging using the same weights [2508.13228].

This architecture indicates that PreSem-Surf couples three layers of modeling: sampling control, geometric field prediction, and semantic supervision. That coupling is central to its reported balance between reconstruction quality and training speed.

## 3. SG-MLP and PR-MLP pre-rendering mechanism

The SG-MLP module is a sampling-guided decoder organized as a two-stage procedure. In the first stage, the PR-MLP performs coarse volumetric estimation over uniformly sampled points $\{x_i\}_{i=1}^N$ in a voxel grid:
$$
\sigma_i = \text{MLP}_\theta\big(\gamma(x_i)\big),
$$
where $\text{MLP}_\theta$ is the lightweight PR-MLP and $\gamma(x_i)$ is NeRF-style positional encoding. This produces a coarse estimate of density across the whole scene and quickly captures large-scale geometry such as walls, floors, and major furniture [2508.13228].

In the second stage, hierarchical progressive sampling is driven by a dynamic threshold:
$$
\tau_{k+1} = \lambda \cdot \frac{1}{N} \sum_{j=1}^{N} \sigma_k(x_j) + (1-\lambda) \cdot \max_{1 \leq j \leq N} \sigma_k(x_j),
$$
with $\lambda \in [0,1]$ balancing mean and maximum density. Points satisfying $\sigma_k(x_d) > \tau_{k+1}$ are retained, and importance sampling in retained regions is defined by
$$
p_{k+1}(x_d) = \frac{\sigma_k(x_d)}{\sum_{j \mid \sigma_k(x_j) > \tau_{k+1} \sigma_k(x_j)}.
$$
Higher-density regions are therefore more likely to be resampled and refined in subsequent layers [2508.13228].

The PR-MLP itself is supervised by a PR loss in a truncated distance region:
$$
\mathcal{L}_{\text{PR} = \frac{1}{|S_{\text{tr}|} \sum_{p \in S_{\text{tr} (D_p - \hat{D}_p)^2,
$$
where $S_{\text{tr}}$ is the set of points in the truncated region, $D_p$ is ground-truth distance, and $\hat{D}_p$ is the PR-MLP prediction. This preconditioning stage is described as providing early rough geometry, helping separate noise from fine details, and reducing the burden on the main SDF MLP [2508.13228].

A plausible implication is that SG-MLP and PR-MLP form a learned allocation policy for sample density: the network does not merely predict geometry, but first predicts where geometry should be queried more intensively. That is the method’s principal efficiency mechanism.

## 4. Progressive semantic modeling

Semantic supervision in PreSem-Surf is not taken from native scene labels in the training set. Instead, DFormer is used to segment RGB images, and each semantic category is mapped to a specific color under the NYU40 label system, producing pseudo-color semantic images that function as semantic targets during training [2508.13228].

PFPSMS operates in two phases. During the first half of training, voxel dimensions are 10× larger than in the fine-precision phase, coarse-grained feature planes and SG-MLP sampling are used, and rendering is performed with NEUS weights based on densities:
$$
\omega_{\text{coarse}(k) = \exp\left(-\sum_{j=1}^{k-1} \sigma_j \Delta z_j\right) \cdot \left(1 - \exp(-\sigma_k \Delta z_k)\right),
$$
with $\sigma_j = \varphi(\text{sdf}_j \cdot \text{inv}_s)$. This stage is described as capturing the global structure and coarse semantic layout while remaining computationally cheap [2508.13228].

During the second half of training, voxel size returns to standard fine resolution, the number of voxels increases substantially, and fine-grained rendering uses ray weights inherited from the coarse stage:
$$
\omega_{\text{fine}(k) = \beta \cdot \omega_{\text{coarse}(k) \cdot \frac{e^{-\sigma_k \Delta z_k}{1 - \exp(-\sigma_{k+1} \Delta z_{k+1})},
$$
where $\beta$ is a scaling factor and $\sigma_k$, $\Delta z_k$ are the fine-phase density and distance terms. The stated purpose of this linkage is to ensure that fine sampling respects coarse structure and to stabilize optimization [2508.13228].

Semantic supervision is introduced through
$$
\mathcal{L}_{\text{sem} = \lambda'_{\text{rgb} \cdot \mathcal{L}'_{\text{rgb} + \lambda'_d \cdot \mathcal{L}'_d,
$$
where $\mathcal{L}'_{\text{rgb}}$ is the color loss between rendered pseudo-color images and DFormer-generated pseudo-color targets, and $\mathcal{L}'_d$ is a depth loss aligned with semantic modeling. The exact forms are described as analogous to RGB and depth losses but applied in pseudo-color space and semantic-guided depth space [2508.13228].

This design suggests that semantic information is used to regularize both occupancy structure and object delineation. The coarse-to-fine schedule also implies that semantic guidance is strongest when determining global scene organization and then progressively refined as geometry becomes more precise.

## 5. Data fusion, optimization, and reconstruction objective

For each scene, PreSem-Surf consumes RGB images, noisy synthetic depth maps, and pseudo-color semantic images derived from DFormer segmentation and NYU40 mapping. These per-pixel signals are mapped to rays in the voxel grid, and SG-MLP selects sample points along the rays [2508.13228].

RGB consistency is enforced by
$$
\mathcal{L}_{\text{rgb} = \frac{1}{N_{\text{rgb} \sum_{m = 1}^{N_{\text{rgb} \mathcal{L}_{\text{rgb}, m},
$$
where $\mathcal{L}_{\text{rgb},m}$ measures the discrepancy between rendered and ground-truth RGB for ray $m$. Depth consistency is enforced by
$$
\mathcal{L}_d = \frac{1}{|\mathcal{R}_d|} \sum_{r \in \mathcal{R}_d} \ell_d^r,
$$
where $\mathcal{R}_d$ is the set of rays with valid depth and $\ell_d^r$ is the per-ray rendered-versus-true depth discrepancy [2508.13228].

The overall objective is given as
$$
\mathcal{L} = \lambda_{\text{SG} \, \mathcal{L}_{\text{SG} + \lambda_{\text{sem} \, \mathcal{L}_{\text{sem}.
$$
The SG-MLP loss expands to
$$
\begin{split}
\mathcal{L}_{\text{SG} = &\ \lambda_{\text{PR} \mathcal{L}_{\text{PR} + \lambda_{\text{rgb} \mathcal{L}_{\text{rgb} + \lambda_d \mathcal{L}_d \\
&+ \lambda_{\text{sdf} \mathcal{L}_{\text{sdf} + \lambda_{\text{eik} \mathcal{L}_{\text{eik} + \lambda_{\text{smooth} \mathcal{L}_{\text{smooth}.
\end{split}
$$
The terms include PR-MLP distance supervision, RGB and depth losses, a truncated-region SDF loss, an FS loss, Eikonal regularization enforcing $\|\nabla \text{SDF}\| \approx 1$, and a smoothness regularizer for near-surface points [2508.13228].

Training proceeds by preprocessing RGB images with DFormer, building voxel grids at multiple resolutions, adding noise and artifacts to synthetic depth to simulate real sensors, using PR-MLP for coarse uniform voxel sampling, then applying SG-MLP thresholding and importance sampling. Optimization uses Adam with learning rate $0.001$ for NeRF/SDF, SG-MLP, and the semantic model; example loss weights are $\lambda_{\text{model} = 5$, $\lambda_{\text{SG} = 4$, and $\lambda_{\text{sem} = 1$ [2508.13228].

After convergence, the SDF MLP defines a continuous signed distance field, and the surface is extracted as the zero level set
$$
\{ \mathbf{x} \mid \text{SDF}(\mathbf{x}) = 0 \}.
$$
A voxel grid is evaluated and marching cubes or a similar procedure is applied to extract a triangle mesh. The stated role of $\mathcal{L}_{\text{eik}}$ and $\mathcal{L}_{\text{smooth}}$ is to keep the field well behaved and the mesh smooth while preserving detail [2508.13228].

## 6. Experimental results and empirical behavior

The experiments use seven scenes from the Synthetic RGB-D dataset of Azinovic et al., with added depth noise and artifacts for realism. Evaluation is reported with C-L1, NC, F-score, IoU, Acc, and Comp. C-L1 is defined as the Chamfer L1 distance between reconstructed and ground-truth surfaces, NC as normal consistency, F-score as the harmonic mean of precision and recall at a distance threshold, IoU as occupied-voxel intersection-over-union, Acc as distance-based geometry accuracy, and Comp as completeness [2508.13228].

Across all scenes, the reported averages are as follows:

| Method | C-L1 | F-score | IoU |
|---|---:|---:|---:|
| Co-SLAM | 0.0573 | 0.8827 | 0.5459 |
| NeRF-SLAM Benchmark | 0.0836 | 0.8375 | 0.4776 |
| Neural RGBD | 0.0261 | 0.9314 | 0.5938 |
| GO-Surf | 0.0264 | 0.9329 | 0.5850 |
| PreSem-Surf | **0.0236** | **0.9440** | **0.6389** |

PreSem-Surf is further reported with NC $0.9132$, very close to the best $0.9138$, Acc $0.0204$, and Comp $0.0250$, both described as competitive though slightly above the best values [2508.13228].

The ablation study isolates the impact of semantics and SG-MLP. Under the reported setting, GO-Surf yields C-L1 $0.0398$, NC $0.9209$, F-score $0.9059$, and IoU $0.5358$. Removing the semantic model gives C-L1 $0.0409$, F-score $0.9062$, and IoU $0.5370$, which is described as showing that semantics improve some aspects but may slightly hurt others. Removing SG-MLP yields C-L1 $0.0462$ and IoU $0.5176$, indicating that SG-MLP is critical for global structure and sampling efficiency. The full PreSem-Surf variant in that evaluation setting yields C-L1 $0.0229$, NC $0.9193$, F-score $0.9186$, and IoU $0.5629$, which is described as highlighting SG-MLP and semantics synergy [2508.13228].

Qualitatively, GO-Surf is described as smooth overall but prone to fragmentation and bloating artifacts, Neural RGBD as smoother with fewer fragments but susceptible to misalignment and misjudgment, and PreSem-Surf as offering a better balance of smoothness and detail, with continuous surfaces, sharper object boundaries, fewer holes and fragments, and better alignment with depth and semantics [2508.13228].

A common misconception would be to treat the semantic component as the sole driver of the gains. The ablation results indicate a more nuanced picture: SG-MLP has a strong impact on geometric accuracy, while semantic modeling improves overall reconstruction but interacts with some metrics in non-uniform ways [2508.13228].

## 7. Computational profile, relation to SURF, and limitations

A parameters analysis on three scenes reports the following scaling behavior. For Morning Apartment, with dimensions $3.8 \times 2.9 \times 4.8$ m and voxel grid $129 \times 97 \times 161$, runtime is approximately 36 minutes with model size 82 MB and 21.5M parameters. For ScanNet Scene 0000, with dimensions $9.6 \times 9.6 \times 3.8$ m and voxel grid $321 \times 321 \times 129$, runtime is approximately 73 minutes with model size 535 MB and 140.4M parameters. For ScanNet Scene 0012, with dimensions $6.7 \times 6.7 \times 3.8$ m and voxel grid $225 \times 225 \times 229$, runtime is approximately 59 minutes with model size 263 MB and 69.1M parameters [2508.13228].

These figures are interpreted in the source as evidence that the method is significantly faster than naive NeRF training for similar resolutions because of coarse pre-rendering, hierarchical SG-MLP sampling, and the PFPSMS coarse-to-fine schedule. At the same time, memory usage is reported to increase rapidly with scene size, which is described as typical of voxel-based representations [2508.13228].

The method’s limitations are also explicit. Memory and time cost grow quickly with scene scale due to dense voxel grids. Performance depends on DFormer-produced semantic labels, so inaccurate or biased segmentation can slightly degrade some metrics. The synthetic dataset includes noisy depth to approximate real sensors, but real-world factors such as lighting variation, motion, and severe depth artifacts may pose additional challenges. Very large scenes, severely missing depth, and some fine normal or completeness cases remain difficult [2508.13228].

Despite the “Surf” in its name, PreSem-Surf is not a SURF keypoint descriptor method. The only direct connection available in the provided record is terminological: one paper on SURF implementation quality emphasizes that any SURF-based variant depends heavily on detector and descriptor stability, interpolation, border handling, and implementation details, with the broader conclusion that reported gains can otherwise be confounded by implementation artifacts [1202.0492]. This suggests that “Surf” in PreSem-Surf should not be conflated with classical Speeded Up Robust Features unless an explicit algorithmic dependence is stated. In the material summarized here, PreSem-Surf is instead an RGB-D NeRF/SDF surface reconstruction framework with semantic modeling [2508.13228].

Within related work, PreSem-Surf is positioned at the intersection of NeRF/SDF-based RGB-D reconstruction, semantic NeRF or semantic SLAM, and voxel-based NeRF acceleration. Its stated novelties are the SG-MLP plus PR-MLP pre-rendering mechanism, the PFPSMS progressive semantic strategy, and the joint integration of RGB, noisy depth, and pseudo-semantic labels into a single reconstruction objective [2508.13228].

Source: https://www.emergentmind.com/topics/presem-surf