TopoSurfel Framework: Geometric Closure
- The TopoSurfel framework closes the geometric gap between 3D Gaussian Splatting primitives and continuous surface representations by dynamically constructing a differentiable proxy mesh
- The proxy mesh uses weighted oriented-point sampling, differentiable Poisson reconstruction, and differentiable marching cubes.
- During optimization, the differentiable proxy mesh provides geometric supervision, improving surfel coherence to surface reconstruction, and is regenerated dynamically during training.
TopoSurfel is a framework for closing the geometric gap between unstructured 3D Gaussian Splatting primitives and continuous surface representations. It represents a scene with anisotropic Gaussian surfels while dynamically constructing a differentiable proxy mesh through weighted oriented-point sampling, differentiable Poisson surface reconstruction, and differentiable marching cubes. The proxy mesh supplies training-time geometric supervision and guides surfel normal orientation, pruning, densification, and initialization, while Gaussian surfels retain their role in differentiable appearance rendering (Fan et al., 21 Aug 2026).
1. Problem formulation and design objectives
3D Gaussian Splatting represents scenes with discrete, unstructured primitives that can reproduce training views while retaining ambiguous or incorrect geometry. In weakly observed regions, including textureless, reflective, occluded, or sparsely viewed areas, photometric supervision may be insufficient to determine a coherent surface. Typical consequences include floaters, holes, fragmented geometry, and structural ambiguity.
TopoSurfel addresses this problem by making a continuous mesh an active geometric prior during optimization rather than extracting a mesh only after Gaussian training. Its principal design loop is:
The proxy mesh is dynamically recomputed and is not an independently optimized parameter set. Unlike approaches that introduce auxiliary neural networks or additional trainable variables for mesh extraction, TopoSurfel uses a parameter-free geometric conversion whose computations are differentiable.
The framework is related to several preceding surface and surfel paradigms. Piecewise-planar surfels can encode local surface geometry more plausibly than independent pixel depths (Trabes et al., 2019). Normal-aware tangent-plane projection and primal-dual operators provide a geometry-sensitive treatment of voxel-face surfels (0802.1617). Surface extraction from cubical complexes demonstrates how topology and connectivity can be preserved during reconstruction (Algarni et al., 2017). These works motivate aspects of local geometry, explicit surface organization, and topology-aware processing, but TopoSurfel itself is a Gaussian–mesh co-evolution framework.
2. Gaussian-surfel representation
Each Gaussian is parameterized by a center, rotation, scale, opacity, and spherical-harmonic appearance:
The covariance is
where is the Gaussian center, its rotation, its scale vector, its opacity, and its view-dependent appearance coefficients.
The trainable variables are the Gaussian centers, rotations, scales, opacities, and spherical-harmonic parameters. Scale regularization continuously compresses the shortest Gaussian axis, converting volumetric Gaussians into approximately planar Gaussian surfels. The direction of this shortest axis defines an intrinsic surfel normal, denoted , up to sign ambiguity. Flattening is therefore not merely a geometric post-processing step: it creates primitives suitable for oriented-point surface reconstruction.
Each retained surfel is expanded into a finite patch during proxy-mesh construction. For an opacity-filtered surfel set
the default opacity threshold is
0
Five oriented samples are generated for each surfel: one center and four tangent-plane offsets,
1
with
2
The center weight is
3
and each offset receives
4
All five samples inherit the same oriented normal. The resulting point set is
5
The five-point construction represents patch extent more effectively than center-only sampling. Area- and KNN-density-weighted alternatives were evaluated, but five-point opacity weighting was retained as a quality–efficiency compromise. The area and local density quantities considered in the supplementary analysis are
6
and
7
Five-point KNN-density weighting slightly improves the TNT F1 score but increases runtime.
3. Two-stage optimization and initialization
TopoSurfel uses a warm-up stage followed by surfel–mesh co-evolution.
Warm-up stage
The first stage performs 10,000 iterations of local Gaussian/surfel optimization. It combines photometric reconstruction, shortest-axis scale regularization, multi-view geometric and photometric consistency, and standard Gaussian densification and pruning.
The warm-up is intended to prevent the initial sparse Structure-from-Motion points from becoming an unstructured Gaussian configuration. Direct Poisson reconstruction from poorly organized SfM-initialized Gaussians can produce a distorted scalar field and unstable topology.
During warm-up, Gaussian opacities are reset every 3,000 iterations. Densification and pruning begin at iteration 500 and are performed every 100 iterations until iteration 9,000.
TSDF-based spatial re-initialization
After warm-up, rendered depth maps are fused by conventional TSDF fusion to produce a coarse mesh scaffold. One surfel is initialized for each TSDF triangle. Its center is the triangle centroid, its local 8-axis is the triangle normal, and two orthogonal in-plane directions define its local frame. Tangential scales are estimated from the triangle extent, the normal scale is made very small, and spherical-harmonic appearance is initialized from the colors of the triangle’s three vertices.
This procedure establishes an approximately one-to-one discrete approximation of the coarse mesh.
Hybrid initialization for large scenes
Finite-resolution TSDF fusion and depth truncation may cover only a central object, leaving background regions unsupported. TopoSurfel therefore computes the nearest-surface distance 9 from every surfel to the TSDF mesh. A surfel is considered mesh-covered if
0
where 1 is the scene scale and the default coefficient is
2
Mesh-covered surfels are reinitialized from TSDF triangles, whereas uncovered background surfels retain their warm-up parameters. This spatially aware hybrid rule preserves a geometric scaffold on the principal object without removing background primitives needed for large-scale reconstruction.
The second, surfel–mesh co-evolution stage lasts 10,000 iterations. Proxy-mesh supervision is activated after 500 iterations. Densification and pruning continue until iteration 7,000 of this stage, with nearest-surface correspondences updated every 100 iterations.
4. Differentiable proxy-mesh construction
At each optimization iteration, the current surfels are converted into a continuous proxy mesh through weighted oriented-point sampling, differentiable Poisson field generation, and differentiable marching cubes.
Differentiable Poisson reconstruction
The oriented samples are scattered onto a regular voxel grid through trilinear interpolation, producing a discrete vector field 3. DPSR then solves a Poisson equation in the frequency domain. With
4
the scalar field is computed as
5
Here 6 is the spatial-frequency coordinate, 7 is the frequency-domain smoothing kernel, and 8 is the scalar indicator field.
Before DPSR, points are normalized into the canonical cube 9 using the current scene center and half-extent. The default training configuration uses a 0 grid. The transformation from surfel parameters to samples, grid scattering, FFT/IFFT operations, and scalar-field processing is differentiable.
Differentiable marching cubes
The scalar field is converted into an explicit mesh with DiffMC. The implementation uses a fixed zero iso-surface level, while the general edge-interpolation equation is
1
where 2 and 3 are voxel-edge endpoints and 4 are their scalar values.
The resulting proxy mesh is
5
with vertices 6 and faces 7. Vertex positions are continuous functions of the scalar field, although marching-cubes cell-configuration decisions remain discrete. The mesh branch is therefore differentiable and non-trainable: it has no learned parameters, but its computations permit gradient propagation.
Mesh rendering and losses
The proxy mesh is rendered with nvdiffrast to produce mesh depth and normal images,
8
The surfels are rendered to produce
9
On valid overlapping pixels, the mesh-depth and mesh-normal losses are
0
and
1
The logarithmic depth comparison is relative rather than purely absolute and reduces the influence of large depth magnitudes.
The total objective is
2
Here 3 is the standard 3DGS photometric loss combining 4 and D-SSIM, 5 is shortest-axis scale regularization, 6 is multi-view consistency, and 7 activates mesh supervision after warm-up.
The multi-view loss is
8
Gradients propagate from mesh depth and normals through the proxy mesh, scalar field, oriented-point samples, and surfel parameters.
5. Mesh-guided surfel evolution
The proxy mesh directly controls surfel orientation, pruning, and densification. It is therefore both a differentiable loss target and a discrete geometric guide for maintaining the surfel population.
Normal alignment
The shortest-axis normal of a Gaussian has an inherent sign ambiguity. Incorrect normal orientations corrupt the oriented-point cloud and can destabilize Poisson reconstruction. TopoSurfel resolves this ambiguity in two stages.
First, it accumulates camera-view directions whenever a surfel is visible:
9
For each surfel, the nearest mesh face and its Euclidean distance 0 are found. Let 1 be the nearest face normal and 2 the in-plane surfel scale. The mesh is trusted when
3
and
4
The reported defaults are
5
When both conditions hold, the mesh normal is used as the reference direction:
6
Otherwise, the accumulated view direction is used:
7
The normal sign is then selected by
8
yielding
9
This operation changes only the normal sign, not surfel geometry. Mesh-guided orientation is more reliable near boundaries and under sparse viewpoints than view-only flipping, while random flipping produces substantially worse results.
Geometry-aware pruning
A surfel is removed when it is simultaneously far from the proxy mesh and has low opacity:
0
and
1
The reported pruning threshold is
2
This criterion suppresses detached low-opacity floaters that opacity-only pruning may retain.
Geometry-aware densification
The proxy mesh also identifies under-covered surface regions. If no surfel is the nearest neighbor of a mesh face, that face is treated as a coverage gap. A new surfel is inserted at the face centroid and inherits the face normal, tangent frame, tangential scales, small normal scale, and spherical-harmonic appearance copied from the nearest existing surfel.
This local rule fills holes without globally increasing Gaussian density. Densification and pruning use nearest-surface correspondences updated every 100 iterations. Opacity is not reset during the second stage because opacity determines oriented-point weights.
6. Computational characteristics, evaluation, and limitations
Trainable and dynamically computed quantities
Trainable quantities comprise the standard Gaussian parameters: positions, rotations, scales, opacities, and spherical-harmonic appearance coefficients.
Dynamically computed quantities include the opacity-filtered surfel set, five oriented samples per surfel, sample weights, oriented normals, the voxel vector field, the DPSR scalar field, the DiffMC proxy mesh, mesh depth and normal maps, nearest mesh faces and distances, and mesh-guided pruning and densification decisions.
The oriented-point construction, trilinear scattering, DPSR, FFT/IFFT operations, DiffMC interpolation, KNN nearest-surface queries, TSDF initialization, final TSDF extraction, and nvdiffrast rendering are non-trainable computations. The mesh is therefore a differentiable computational object rather than a second independently optimized representation.
Memory and runtime trade-offs
DPSR uses a dense volumetric grid whose memory cost grows cubically with resolution. In a stress test on an RTX 3090 with 24 GB of memory, DTU grids at resolutions 3, 4, and 5 used 4.2, 8.5, and 17.5 GB, respectively. DTU Chamfer distance improved from 0.54 to 0.51 to 0.50. On TNT, 6 improved F1 from 0.45 at 7 to 0.52, while 8 exceeded memory capacity for three TNT scenes.
Consequently, 9 is used as a practical training-time proxy resolution. Higher resolutions can preserve thinner structures but are expensive for large scenes. TSDF fusion is used for final mesh export because gradients are no longer required and TSDF is more memory efficient.
The main reported configuration uses one NVIDIA RTX 3090, 10,000 warm-up iterations, 10,000 mesh-guided iterations, a 0 DPSR grid, PyTorch, and nvdiffrast. On the DTU efficiency comparison, TopoSurfel uses approximately 8.5 GB peak GPU memory and 37 minutes of training, with 0.14 million Gaussians, 284 FPS Gaussian rendering, and a final mesh containing 1.6 million vertices with an 83.9 MB representation.
Geometric reconstruction
The evaluation covers DTU, Tanks and Temples, Mip-NeRF 360, and NeRF-Synthetic. Baselines include NeRF-based methods, Gaussian methods such as 2DGS, PGSR, GOF, and QGS, explicit methods such as IMLS-Splatting and GeoSVR, and Gaussian–mesh hybrids including MILo and MeshSplatting.
On DTU, TopoSurfel reports an average Chamfer Distance of 0.51, compared with 0.53 for PGSR, 0.54 for QGS, 0.68 for MILo, and 0.76 for 2DGS. GeoSVR reports 0.47 but uses an additional monocular depth prior.
On Tanks and Temples, TopoSurfel reports a mean F1 score of 0.52, compared with 0.49 for MILo, 0.50 for Neuralangelo, PGSR, and QGS, and 0.40 for RaDe-GS. The reported qualitative improvements include fewer floaters, cleaner boundaries, fewer holes, and more coherent large-scale surfaces.
Novel-view synthesis
On Mip-NeRF 360, the reported averages are:
For mesh-based novel-view synthesis on NeRF-Synthetic, the reconstructed mesh is UV-unwrapped and a 1 texture map is optimized for 3,000 iterations with fixed geometry. TopoSurfel reports
2
with 1.7 million mesh vertices.
The mesh is especially relevant to mesh-based rendering because texture optimization cannot repair holes or disconnected fragments in the geometry. More complete surface topology therefore improves the final textured-mesh representation even when Gaussian image rendering remains the primary training objective.
Ablation results
The TNT ablation reports:
| Variant | F1 | PSNR |
|---|---|---|
| Without mesh loop | 0.42 | 26.75 |
| Without warm-up | 0.39 | 26.46 |
| View-statistics normal flips | 0.46 | 26.91 |
| Random normal flips | 0.37 | 26.20 |
| Without hybrid initialization | 0.21 | 20.59 |
| Without geometry-aware density control | 0.48 | 27.03 |
| Full TopoSurfel | 0.52 | 26.83 |
The results associate the warm-up with stable initial iso-surfacing, the mesh loop with improved geometry, mesh-guided normal alignment with coherent Poisson input, hybrid initialization with large-scale background preservation, and geometry-aware density control with floater suppression and hole filling.
The supplementary comparisons report F1 0.43 and 98 minutes for DPSR final extraction, F1 0.39 and 137 minutes for FlexiCubes extraction, F1 0.53 and 142 minutes for five-point KNN-density sampling, and F1 0.52 and 98 minutes for five-point opacity weighting.
Limitations
TopoSurfel retains several limitations of Gaussian-based reconstruction:
- Dense-grid scalability: DPSR memory grows cubically with grid resolution, restricting high-resolution proxy meshes in large scenes.
- Micro-geometry loss: the default 3 training proxy may smooth thin structures and small details.
- Reflective and transparent surfaces: view-dependent appearance can be explained without recovering physically correct geometry.
- Initialization dependence: the warm-up TSDF scaffold stabilizes the process, but poor initial depth can still affect reconstruction.
- Discrete topology decisions: DiffMC provides differentiable vertex positioning, but marching-cubes configuration changes are not fully smooth optimization variables.
- Additional computational cost: regenerating and rendering a proxy mesh during training increases memory and runtime relative to ordinary Gaussian optimization.
The framework also does not by itself constitute a complete explicit topological data structure. It dynamically constructs a mesh and uses mesh–surfel correspondences, but the supplied formulation does not define persistent homology, a half-edge topology, explicit topological event handling, or guaranteed topology preservation under arbitrary surfel edits. Topological surface representations based on cubical complexes provide stronger explicit connectivity tests, including edge-incidence and cyclic vertex-link conditions (DÃaz et al., 2017), while persistence-guided reconstruction supplies a separate mechanism for imposing prescribed homological structure (Brüel-Gabrielsson et al., 2018).
TopoSurfel’s principal contribution is consequently the integration of Gaussian rendering with a differentiable geometric scaffold: the proxy mesh affects surfel optimization during training, normal orientation, geometry-aware population control, and final surface completeness. It does not replace Gaussian Splatting with a mesh; rather, it couples the two representations so that Gaussian surfels retain appearance flexibility while the dynamically reconstructed mesh supplies global surface organization.