Papers
Topics
Authors
Recent
Search
2000 character limit reached

TopoSurfel Framework: Geometric Closure

Updated 26 August 2026
  • The TopoSurfel framework closes the geometric gap between 3D Gaussian Splatting primitives and continuous surface representations by dynamically constructing a differentiable proxy mesh
  • The proxy mesh uses weighted oriented-point sampling, differentiable Poisson reconstruction, and differentiable marching cubes.
  • During optimization, the differentiable proxy mesh provides geometric supervision, improving surfel coherence to surface reconstruction, and is regenerated dynamically during training.

TopoSurfel is a framework for closing the geometric gap between unstructured 3D Gaussian Splatting primitives and continuous surface representations. It represents a scene with anisotropic Gaussian surfels while dynamically constructing a differentiable proxy mesh through weighted oriented-point sampling, differentiable Poisson surface reconstruction, and differentiable marching cubes. The proxy mesh supplies training-time geometric supervision and guides surfel normal orientation, pruning, densification, and initialization, while Gaussian surfels retain their role in differentiable appearance rendering (Fan et al., 21 Aug 2026).

1. Problem formulation and design objectives

3D Gaussian Splatting represents scenes with discrete, unstructured primitives that can reproduce training views while retaining ambiguous or incorrect geometry. In weakly observed regions, including textureless, reflective, occluded, or sparsely viewed areas, photometric supervision may be insufficient to determine a coherent surface. Typical consequences include floaters, holes, fragmented geometry, and structural ambiguity.

TopoSurfel addresses this problem by making a continuous mesh an active geometric prior during optimization rather than extracting a mesh only after Gaussian training. Its principal design loop is:

Gaussian surfels→oriented points→Poisson field→proxy mesh→mesh supervision and surfel evolution→updated surfels.\text{Gaussian surfels} \rightarrow \text{oriented points} \rightarrow \text{Poisson field} \rightarrow \text{proxy mesh} \rightarrow \text{mesh supervision and surfel evolution} \rightarrow \text{updated surfels}.

The proxy mesh is dynamically recomputed and is not an independently optimized parameter set. Unlike approaches that introduce auxiliary neural networks or additional trainable variables for mesh extraction, TopoSurfel uses a parameter-free geometric conversion whose computations are differentiable.

The framework is related to several preceding surface and surfel paradigms. Piecewise-planar surfels can encode local surface geometry more plausibly than independent pixel depths (Trabes et al., 2019). Normal-aware tangent-plane projection and primal-dual operators provide a geometry-sensitive treatment of voxel-face surfels (0802.1617). Surface extraction from cubical complexes demonstrates how topology and connectivity can be preserved during reconstruction (Algarni et al., 2017). These works motivate aspects of local geometry, explicit surface organization, and topology-aware processing, but TopoSurfel itself is a Gaussian–mesh co-evolution framework.

2. Gaussian-surfel representation

Each Gaussian is parameterized by a center, rotation, scale, opacity, and spherical-harmonic appearance:

Gi=(μi,Ri,Si,αi,SHi).\mathcal{G}_i = \left( \boldsymbol{\mu}_i, R_i, S_i, \alpha_i, \mathrm{SH}_i \right).

The covariance is

Σi=RiSiSiTRiT,\Sigma_i=R_iS_iS_i^{\mathsf T}R_i^{\mathsf T},

where μi\boldsymbol{\mu}_i is the Gaussian center, RiR_i its rotation, SiS_i its scale vector, αi\alpha_i its opacity, and SHi\mathrm{SH}_i its view-dependent appearance coefficients.

The trainable variables are the Gaussian centers, rotations, scales, opacities, and spherical-harmonic parameters. Scale regularization continuously compresses the shortest Gaussian axis, converting volumetric Gaussians into approximately planar Gaussian surfels. The direction of this shortest axis defines an intrinsic surfel normal, denoted ni\mathbf n_i, up to sign ambiguity. Flattening is therefore not merely a geometric post-processing step: it creates primitives suitable for oriented-point surface reconstruction.

Each retained surfel is expanded into a finite patch during proxy-mesh construction. For an opacity-filtered surfel set

S={si∣αi>τopac},\mathcal S= \{s_i\mid \alpha_i>\tau_{\mathrm{opac}}\},

the default opacity threshold is

Gi=(μi,Ri,Si,αi,SHi).\mathcal{G}_i = \left( \boldsymbol{\mu}_i, R_i, S_i, \alpha_i, \mathrm{SH}_i \right).0

Five oriented samples are generated for each surfel: one center and four tangent-plane offsets,

Gi=(μi,Ri,Si,αi,SHi).\mathcal{G}_i = \left( \boldsymbol{\mu}_i, R_i, S_i, \alpha_i, \mathrm{SH}_i \right).1

with

Gi=(μi,Ri,Si,αi,SHi).\mathcal{G}_i = \left( \boldsymbol{\mu}_i, R_i, S_i, \alpha_i, \mathrm{SH}_i \right).2

The center weight is

Gi=(μi,Ri,Si,αi,SHi).\mathcal{G}_i = \left( \boldsymbol{\mu}_i, R_i, S_i, \alpha_i, \mathrm{SH}_i \right).3

and each offset receives

Gi=(μi,Ri,Si,αi,SHi).\mathcal{G}_i = \left( \boldsymbol{\mu}_i, R_i, S_i, \alpha_i, \mathrm{SH}_i \right).4

All five samples inherit the same oriented normal. The resulting point set is

Gi=(μi,Ri,Si,αi,SHi).\mathcal{G}_i = \left( \boldsymbol{\mu}_i, R_i, S_i, \alpha_i, \mathrm{SH}_i \right).5

The five-point construction represents patch extent more effectively than center-only sampling. Area- and KNN-density-weighted alternatives were evaluated, but five-point opacity weighting was retained as a quality–efficiency compromise. The area and local density quantities considered in the supplementary analysis are

Gi=(μi,Ri,Si,αi,SHi).\mathcal{G}_i = \left( \boldsymbol{\mu}_i, R_i, S_i, \alpha_i, \mathrm{SH}_i \right).6

and

Gi=(μi,Ri,Si,αi,SHi).\mathcal{G}_i = \left( \boldsymbol{\mu}_i, R_i, S_i, \alpha_i, \mathrm{SH}_i \right).7

Five-point KNN-density weighting slightly improves the TNT F1 score but increases runtime.

3. Two-stage optimization and initialization

TopoSurfel uses a warm-up stage followed by surfel–mesh co-evolution.

Warm-up stage

The first stage performs 10,000 iterations of local Gaussian/surfel optimization. It combines photometric reconstruction, shortest-axis scale regularization, multi-view geometric and photometric consistency, and standard Gaussian densification and pruning.

The warm-up is intended to prevent the initial sparse Structure-from-Motion points from becoming an unstructured Gaussian configuration. Direct Poisson reconstruction from poorly organized SfM-initialized Gaussians can produce a distorted scalar field and unstable topology.

During warm-up, Gaussian opacities are reset every 3,000 iterations. Densification and pruning begin at iteration 500 and are performed every 100 iterations until iteration 9,000.

TSDF-based spatial re-initialization

After warm-up, rendered depth maps are fused by conventional TSDF fusion to produce a coarse mesh scaffold. One surfel is initialized for each TSDF triangle. Its center is the triangle centroid, its local Gi=(μi,Ri,Si,αi,SHi).\mathcal{G}_i = \left( \boldsymbol{\mu}_i, R_i, S_i, \alpha_i, \mathrm{SH}_i \right).8-axis is the triangle normal, and two orthogonal in-plane directions define its local frame. Tangential scales are estimated from the triangle extent, the normal scale is made very small, and spherical-harmonic appearance is initialized from the colors of the triangle’s three vertices.

This procedure establishes an approximately one-to-one discrete approximation of the coarse mesh.

Hybrid initialization for large scenes

Finite-resolution TSDF fusion and depth truncation may cover only a central object, leaving background regions unsupported. TopoSurfel therefore computes the nearest-surface distance Gi=(μi,Ri,Si,αi,SHi).\mathcal{G}_i = \left( \boldsymbol{\mu}_i, R_i, S_i, \alpha_i, \mathrm{SH}_i \right).9 from every surfel to the TSDF mesh. A surfel is considered mesh-covered if

Σi=RiSiSiTRiT,\Sigma_i=R_iS_iS_i^{\mathsf T}R_i^{\mathsf T},0

where Σi=RiSiSiTRiT,\Sigma_i=R_iS_iS_i^{\mathsf T}R_i^{\mathsf T},1 is the scene scale and the default coefficient is

Σi=RiSiSiTRiT,\Sigma_i=R_iS_iS_i^{\mathsf T}R_i^{\mathsf T},2

Mesh-covered surfels are reinitialized from TSDF triangles, whereas uncovered background surfels retain their warm-up parameters. This spatially aware hybrid rule preserves a geometric scaffold on the principal object without removing background primitives needed for large-scale reconstruction.

The second, surfel–mesh co-evolution stage lasts 10,000 iterations. Proxy-mesh supervision is activated after 500 iterations. Densification and pruning continue until iteration 7,000 of this stage, with nearest-surface correspondences updated every 100 iterations.

4. Differentiable proxy-mesh construction

At each optimization iteration, the current surfels are converted into a continuous proxy mesh through weighted oriented-point sampling, differentiable Poisson field generation, and differentiable marching cubes.

Differentiable Poisson reconstruction

The oriented samples are scattered onto a regular voxel grid through trilinear interpolation, producing a discrete vector field Σi=RiSiSiTRiT,\Sigma_i=R_iS_iS_i^{\mathsf T}R_i^{\mathsf T},3. DPSR then solves a Poisson equation in the frequency domain. With

Σi=RiSiSiTRiT,\Sigma_i=R_iS_iS_i^{\mathsf T}R_i^{\mathsf T},4

the scalar field is computed as

Σi=RiSiSiTRiT,\Sigma_i=R_iS_iS_i^{\mathsf T}R_i^{\mathsf T},5

Here Σi=RiSiSiTRiT,\Sigma_i=R_iS_iS_i^{\mathsf T}R_i^{\mathsf T},6 is the spatial-frequency coordinate, Σi=RiSiSiTRiT,\Sigma_i=R_iS_iS_i^{\mathsf T}R_i^{\mathsf T},7 is the frequency-domain smoothing kernel, and Σi=RiSiSiTRiT,\Sigma_i=R_iS_iS_i^{\mathsf T}R_i^{\mathsf T},8 is the scalar indicator field.

Before DPSR, points are normalized into the canonical cube Σi=RiSiSiTRiT,\Sigma_i=R_iS_iS_i^{\mathsf T}R_i^{\mathsf T},9 using the current scene center and half-extent. The default training configuration uses a μi\boldsymbol{\mu}_i0 grid. The transformation from surfel parameters to samples, grid scattering, FFT/IFFT operations, and scalar-field processing is differentiable.

Differentiable marching cubes

The scalar field is converted into an explicit mesh with DiffMC. The implementation uses a fixed zero iso-surface level, while the general edge-interpolation equation is

μi\boldsymbol{\mu}_i1

where μi\boldsymbol{\mu}_i2 and μi\boldsymbol{\mu}_i3 are voxel-edge endpoints and μi\boldsymbol{\mu}_i4 are their scalar values.

The resulting proxy mesh is

μi\boldsymbol{\mu}_i5

with vertices μi\boldsymbol{\mu}_i6 and faces μi\boldsymbol{\mu}_i7. Vertex positions are continuous functions of the scalar field, although marching-cubes cell-configuration decisions remain discrete. The mesh branch is therefore differentiable and non-trainable: it has no learned parameters, but its computations permit gradient propagation.

Mesh rendering and losses

The proxy mesh is rendered with nvdiffrast to produce mesh depth and normal images,

μi\boldsymbol{\mu}_i8

The surfels are rendered to produce

μi\boldsymbol{\mu}_i9

On valid overlapping pixels, the mesh-depth and mesh-normal losses are

RiR_i0

and

RiR_i1

The logarithmic depth comparison is relative rather than purely absolute and reduces the influence of large depth magnitudes.

The total objective is

RiR_i2

Here RiR_i3 is the standard 3DGS photometric loss combining RiR_i4 and D-SSIM, RiR_i5 is shortest-axis scale regularization, RiR_i6 is multi-view consistency, and RiR_i7 activates mesh supervision after warm-up.

The multi-view loss is

RiR_i8

Gradients propagate from mesh depth and normals through the proxy mesh, scalar field, oriented-point samples, and surfel parameters.

5. Mesh-guided surfel evolution

The proxy mesh directly controls surfel orientation, pruning, and densification. It is therefore both a differentiable loss target and a discrete geometric guide for maintaining the surfel population.

Normal alignment

The shortest-axis normal of a Gaussian has an inherent sign ambiguity. Incorrect normal orientations corrupt the oriented-point cloud and can destabilize Poisson reconstruction. TopoSurfel resolves this ambiguity in two stages.

First, it accumulates camera-view directions whenever a surfel is visible:

RiR_i9

For each surfel, the nearest mesh face and its Euclidean distance SiS_i0 are found. Let SiS_i1 be the nearest face normal and SiS_i2 the in-plane surfel scale. The mesh is trusted when

SiS_i3

and

SiS_i4

The reported defaults are

SiS_i5

When both conditions hold, the mesh normal is used as the reference direction:

SiS_i6

Otherwise, the accumulated view direction is used:

SiS_i7

The normal sign is then selected by

SiS_i8

yielding

SiS_i9

This operation changes only the normal sign, not surfel geometry. Mesh-guided orientation is more reliable near boundaries and under sparse viewpoints than view-only flipping, while random flipping produces substantially worse results.

Geometry-aware pruning

A surfel is removed when it is simultaneously far from the proxy mesh and has low opacity:

αi\alpha_i0

and

αi\alpha_i1

The reported pruning threshold is

αi\alpha_i2

This criterion suppresses detached low-opacity floaters that opacity-only pruning may retain.

Geometry-aware densification

The proxy mesh also identifies under-covered surface regions. If no surfel is the nearest neighbor of a mesh face, that face is treated as a coverage gap. A new surfel is inserted at the face centroid and inherits the face normal, tangent frame, tangential scales, small normal scale, and spherical-harmonic appearance copied from the nearest existing surfel.

This local rule fills holes without globally increasing Gaussian density. Densification and pruning use nearest-surface correspondences updated every 100 iterations. Opacity is not reset during the second stage because opacity determines oriented-point weights.

6. Computational characteristics, evaluation, and limitations

Trainable and dynamically computed quantities

Trainable quantities comprise the standard Gaussian parameters: positions, rotations, scales, opacities, and spherical-harmonic appearance coefficients.

Dynamically computed quantities include the opacity-filtered surfel set, five oriented samples per surfel, sample weights, oriented normals, the voxel vector field, the DPSR scalar field, the DiffMC proxy mesh, mesh depth and normal maps, nearest mesh faces and distances, and mesh-guided pruning and densification decisions.

The oriented-point construction, trilinear scattering, DPSR, FFT/IFFT operations, DiffMC interpolation, KNN nearest-surface queries, TSDF initialization, final TSDF extraction, and nvdiffrast rendering are non-trainable computations. The mesh is therefore a differentiable computational object rather than a second independently optimized representation.

Memory and runtime trade-offs

DPSR uses a dense volumetric grid whose memory cost grows cubically with resolution. In a stress test on an RTX 3090 with 24 GB of memory, DTU grids at resolutions αi\alpha_i3, αi\alpha_i4, and αi\alpha_i5 used 4.2, 8.5, and 17.5 GB, respectively. DTU Chamfer distance improved from 0.54 to 0.51 to 0.50. On TNT, αi\alpha_i6 improved F1 from 0.45 at αi\alpha_i7 to 0.52, while αi\alpha_i8 exceeded memory capacity for three TNT scenes.

Consequently, αi\alpha_i9 is used as a practical training-time proxy resolution. Higher resolutions can preserve thinner structures but are expensive for large scenes. TSDF fusion is used for final mesh export because gradients are no longer required and TSDF is more memory efficient.

The main reported configuration uses one NVIDIA RTX 3090, 10,000 warm-up iterations, 10,000 mesh-guided iterations, a SHi\mathrm{SH}_i0 DPSR grid, PyTorch, and nvdiffrast. On the DTU efficiency comparison, TopoSurfel uses approximately 8.5 GB peak GPU memory and 37 minutes of training, with 0.14 million Gaussians, 284 FPS Gaussian rendering, and a final mesh containing 1.6 million vertices with an 83.9 MB representation.

Geometric reconstruction

The evaluation covers DTU, Tanks and Temples, Mip-NeRF 360, and NeRF-Synthetic. Baselines include NeRF-based methods, Gaussian methods such as 2DGS, PGSR, GOF, and QGS, explicit methods such as IMLS-Splatting and GeoSVR, and Gaussian–mesh hybrids including MILo and MeshSplatting.

On DTU, TopoSurfel reports an average Chamfer Distance of 0.51, compared with 0.53 for PGSR, 0.54 for QGS, 0.68 for MILo, and 0.76 for 2DGS. GeoSVR reports 0.47 but uses an additional monocular depth prior.

On Tanks and Temples, TopoSurfel reports a mean F1 score of 0.52, compared with 0.49 for MILo, 0.50 for Neuralangelo, PGSR, and QGS, and 0.40 for RaDe-GS. The reported qualitative improvements include fewer floaters, cleaner boundaries, fewer holes, and more coherent large-scale surfaces.

Novel-view synthesis

On Mip-NeRF 360, the reported averages are:

Scene category PSNR SSIM LPIPS
Indoor 30.47 0.918 0.173
Outdoor 24.18 0.712 0.240

For mesh-based novel-view synthesis on NeRF-Synthetic, the reconstructed mesh is UV-unwrapped and a SHi\mathrm{SH}_i1 texture map is optimized for 3,000 iterations with fixed geometry. TopoSurfel reports

SHi\mathrm{SH}_i2

with 1.7 million mesh vertices.

The mesh is especially relevant to mesh-based rendering because texture optimization cannot repair holes or disconnected fragments in the geometry. More complete surface topology therefore improves the final textured-mesh representation even when Gaussian image rendering remains the primary training objective.

Ablation results

The TNT ablation reports:

Variant F1 PSNR
Without mesh loop 0.42 26.75
Without warm-up 0.39 26.46
View-statistics normal flips 0.46 26.91
Random normal flips 0.37 26.20
Without hybrid initialization 0.21 20.59
Without geometry-aware density control 0.48 27.03
Full TopoSurfel 0.52 26.83

The results associate the warm-up with stable initial iso-surfacing, the mesh loop with improved geometry, mesh-guided normal alignment with coherent Poisson input, hybrid initialization with large-scale background preservation, and geometry-aware density control with floater suppression and hole filling.

The supplementary comparisons report F1 0.43 and 98 minutes for DPSR final extraction, F1 0.39 and 137 minutes for FlexiCubes extraction, F1 0.53 and 142 minutes for five-point KNN-density sampling, and F1 0.52 and 98 minutes for five-point opacity weighting.

Limitations

TopoSurfel retains several limitations of Gaussian-based reconstruction:

  1. Dense-grid scalability: DPSR memory grows cubically with grid resolution, restricting high-resolution proxy meshes in large scenes.
  2. Micro-geometry loss: the default SHi\mathrm{SH}_i3 training proxy may smooth thin structures and small details.
  3. Reflective and transparent surfaces: view-dependent appearance can be explained without recovering physically correct geometry.
  4. Initialization dependence: the warm-up TSDF scaffold stabilizes the process, but poor initial depth can still affect reconstruction.
  5. Discrete topology decisions: DiffMC provides differentiable vertex positioning, but marching-cubes configuration changes are not fully smooth optimization variables.
  6. Additional computational cost: regenerating and rendering a proxy mesh during training increases memory and runtime relative to ordinary Gaussian optimization.

The framework also does not by itself constitute a complete explicit topological data structure. It dynamically constructs a mesh and uses mesh–surfel correspondences, but the supplied formulation does not define persistent homology, a half-edge topology, explicit topological event handling, or guaranteed topology preservation under arbitrary surfel edits. Topological surface representations based on cubical complexes provide stronger explicit connectivity tests, including edge-incidence and cyclic vertex-link conditions (Díaz et al., 2017), while persistence-guided reconstruction supplies a separate mechanism for imposing prescribed homological structure (Brüel-Gabrielsson et al., 2018).

TopoSurfel’s principal contribution is consequently the integration of Gaussian rendering with a differentiable geometric scaffold: the proxy mesh affects surfel optimization during training, normal orientation, geometry-aware population control, and final surface completeness. It does not replace Gaussian Splatting with a mesh; rather, it couples the two representations so that Gaussian surfels retain appearance flexibility while the dynamically reconstructed mesh supplies global surface organization.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to TopoSurfel.