Incremental Online Scene Reconstruction by 3D Gaussian Triangulation
Abstract: Incremental scene reconstruction is essential for real-world applications. Although 3D Gaussian Splatting shows strong potential, most existing approaches require offline conversion of the optimized Gaussians into an intermediate implicit field for explicit mesh extraction, which hinders seamless integration with downstream tasks. To address this limitation, we propose a novel online framework that incrementally reconstructs and updates high-fidelity explicit meshes by directly triangulating a dense geometric Gaussian representation, which supports both high-quality rendering and incremental surface reconstruction. Moreover, we present a direct meshing algorithm that efficiently extracts and updates the mesh from the Gaussian set. To ensure mesh accuracy, we enforce a plane-based pulling constraint that dynamically aligns 3D Gaussian primitives to the approximated local surface. Furthermore, our framework significantly reduces memory and computational overhead during long-sequence processing by dynamically freezing fully optimized historical regions. Experiments on public datasets demonstrate that our method outperforms conventional Gaussian-based methods on both rendering quality and reconstruction accuracy.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
Overview
This paper is about building 3D models of real-world places quickly as a camera moves around, instead of waiting until the end. The authors show a new way to turn a stream of camera images and depth data into a clean 3D surface (a “mesh”) in real time. They do this by treating the scene as lots of tiny, flat, oval stickers placed on surfaces and then smartly connecting those stickers to form triangles, like building a net over the scene.
What are the main questions?
The paper asks:
- How can we rebuild a 3D scene step by step (incrementally) while a camera is moving, without having to restart from scratch each time?
- Can we turn a popular fast rendering method (called 3D Gaussian Splatting) into a tool that also makes accurate, ready-to-use 3D meshes?
- How can we keep the system fast and memory-efficient for long recordings, like moving through an entire building?
How did they do it?
Think of their method like this: as the camera moves and records color images plus depth (how far things are), the system places many tiny, flat ovals on the surfaces it sees. These ovals are called “Gaussians,” but in this paper they are treated like small flat “surfels” (surface pixels). Each surfel knows:
- where it is in 3D,
- which way it’s facing (its “normal”),
- how big it is,
- how see-through it is (opacity),
- and its color.
Once enough of these surfels are in place, the system connects nearby ones to make triangles, forming a mesh (like connecting dots to make a net).
Here are the key ideas, explained with everyday language:
- Dense geometric Gaussians (flat surfels):
- Instead of “puffy” blobs in space, they flatten the Gaussians into thin, flat ovals that stick to surfaces, like stickers. They also make them mostly opaque so they act like real, solid surfaces.
- Plane “pulling” constraint (keeping stickers on the wall):
- If a surfel drifts off the surface (due to noisy depth or movement), a gentle rule pulls it back towards the nearby flat plane of the real surface. Imagine a magnet that keeps the sticker pressed firmly to the wall.
- Normal alignment (face the right way):
- Each surfel should face the same way as its local neighbors if they’re on the same surface (like tiles on a floor all lying flat and aligned).
- Direct triangulation (connect the dots quickly):
- The system picks the trustworthy surfels (opaque and matching depth well), then, around each surfel, it finds neighbors, projects them onto a local flat plane, and connects them in order to form triangles. This avoids heavy, slow methods that require building huge 3D grids.
- Local remeshing and freezing (finish sections and move on):
- After forming a mesh for the currently viewed area, the system tidies up the triangles (splitting long edges, collapsing short ones, smoothing) so they fit nicely with the existing mesh.
- When a region is well-seen and stable, it’s “frozen” — the system stops spending time and memory on it and focuses on new areas. Think of finishing parts of a big puzzle, gluing them down, and then working on the remaining pieces.
- Incremental updates:
- As new frames arrive, only the nearby, newly seen region is updated; the rest stays put. This makes the whole process fast and scalable.
What did they find?
On standard indoor datasets (Replica and ScanNet++), the method:
- Builds more accurate meshes (closer to ground truth and with fewer holes) than several popular baselines, including traditional TSDF fusion, NeRF-style methods, and other Gaussian-splatting approaches.
- Renders high-quality images from new viewpoints (sharp details, fewer artifacts).
- Runs faster and uses less memory, thanks to the direct triangulation and region freezing.
In simple terms: it makes cleaner, more detailed 3D models, faster, and scales better to big scenes.
Why is this important?
- Real-time use: Robots, AR headsets, and drones often need a partial 3D model quickly to make decisions (navigate, avoid obstacles, place virtual objects). Waiting to process everything at the end is too slow.
- Ready-to-use meshes: Many downstream tasks (physics, collision, planning) need meshes, not just pretty renderings. This method gives meshes right away, without heavy offline steps.
- Efficiency: By freezing finished parts, the system avoids running out of memory or slowing down over time.
Limitations and future work
- It currently relies on depth data (RGB-D). If a region was never seen or has no depth, it’s hard to reconstruct.
- The authors suggest exploring reconstruction using only RGB images in the future, which would make the system more flexible.
In short
The paper turns fast “Gaussian splats” into practical, accurate 3D meshes as a camera moves, by using flat, opaque surfels, gentle geometric rules to keep them on surfaces, a quick “connect-the-dots” triangulation, and a smart way to freeze finished areas. It’s faster, more memory-friendly, and makes better models than many existing methods, which is great for real-world applications like AR and robotics.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
Below is a consolidated list of what remains missing, uncertain, or unexplored in the paper, framed as concrete, actionable directions for future work.
- Reliance on depth sensors and known poses: The method assumes RGB-D input with ground-truth camera poses; robustness to pose noise and intrinsic/extrinsic calibration errors is not studied, nor is joint tracking–mapping or loop-closure handling.
- No RGB-only reconstruction: The authors acknowledge that unobserved regions cannot be reconstructed and that the pipeline depends on depth; monocular (RGB-only) reconstruction and depth/pose inference remain open.
- Lack of loop-closure and drift correction with freezing: The freezing strategy does not describe how to unfreeze or globally reconcile meshes if late pose corrections (e.g., loop closure) occur; mechanisms to detect drift and propagate corrections across frozen regions are missing.
- Static-scene assumption: Handling of dynamic or semi-static scenes (moving objects, non-rigid deformation, scene changes) is not addressed; no strategy for motion segmentation, dynamic object suppression, or re-meshing after dynamics is provided.
- Sparse real-world robustness evaluation: Experiments are limited to indoor datasets (Replica, ScanNet++). Performance under challenging real-world conditions (consumer depth noise, rolling shutter, low light, outdoor environments, large-scale unbounded scenes, severe occlusions) is not evaluated.
- Glass/reflective/transparent surfaces: The approach’s behavior on non-Lambertian and semi-transparent materials (where depth is unreliable) is unexplored; no failure analysis or mitigation strategies are proposed.
- Sensitivity to normal estimation: The plane-based pulling and normal consistency rely on accurate oriented point clouds; robustness to erroneous normals and strategies for robust normal estimation (e.g., multi-scale, learned normals) are not analyzed.
- Geometric Gaussian set selection: The hard-threshold selection based on opacity (>0.9) and depth fidelity (τ) is heuristic and scene-dependent; adaptive, multi-view-consistent, or learned selection criteria are not explored beyond a τ ablation.
- Triangulation guarantees: The greedy, angle-based triangulation lacks formal guarantees of watertightness, manifoldness, absence of self-intersections, or topological correctness; no evaluation of topological metrics is provided.
- Sharp features and thin structures: Normal-consistency pruning (e.g., cosine > 0.9) may break connectivity at edges and high-curvature regions; edge-aware neighbor selection or feature-preserving triangulation is not developed or evaluated.
- Over-triangulation/bridging across gaps: Despite mutual visibility and normal checks, failure cases where neighboring surfels connect across depth discontinuities or occlusion boundaries are not analyzed.
- Meshing stability over time: Temporal consistency of the incremental meshing (e.g., flicker, topology churn, vertex drift across updates) is not quantified; methods for temporal regularization or hysteresis in connectivity are absent.
- Local remeshing side effects: The isotropic remeshing (edge splits/collapses, flips, Laplacian smoothing) may blur sharp features; no quantitative assessment of detail loss or edge preservation is provided.
- Interaction between mesh and Gaussian optimization: Triangulation is not integrated into the optimization loop as a differentiable component; mesh-aware losses (e.g., curvature, smoothness, silhouette consistency) are not leveraged.
- Depth rendering model clarity: The piecewise definition for depth via “front opaque Gaussian intersection” (with grazing-angle condition) is underspecified and not validated against ground-truth depth statistics; its failure modes (e.g., grazing angles, multi-surface intersections) are not explored.
- Densification policy: While Gaussians are added based on color/depth errors, a principled schedule for densification (e.g., multi-view confidence, uncertainty, coverage completeness) is not detailed; ablations on densification criteria are missing.
- Parameterization rigidity (s3 ≈ 0): Forcing planar surfels (s3 ≈ 0) simplifies meshing but limits modeling of micro-curvature within a surfel; adaptive s3 or multi-scale anisotropy to better fit curved surfaces is unexplored.
- Handling of transparent Gaussians: The pipeline introduces transparent Gaussians to boost rendering quality but does not study memory/performance trade-offs, their impact on geometric selection, or mechanisms to avoid photometric/geometry conflicts.
- Uncertainty quantification: No per-region or per-triangle uncertainty/confidence is reported; uncertainty-driven selection, freezing, or downstream consumption (e.g., planning) remains open.
- Freezing criteria and premature convergence: The thresholds for observation count and loss convergence (Nobs, εgs) are heuristic; detection of premature freezing and policies for re-activation are not investigated.
- Scalability and out-of-core operation: Although freezing limits growth, there is no out-of-core data management (disk-backed octrees, tiled streaming) or multi-GPU scaling shown for very large scenes.
- Downstream task integration: Despite the motivation to “seamlessly integrate with downstream tasks,” no demonstrations (e.g., AR occlusion handling, robotic planning, grasping) or interfaces for real-time consumption of the incremental mesh are presented.
- Comparative breadth: Direct comparisons with other Gaussian-to-mesh methods (e.g., SuGaR, 2DGS) are limited, especially on large-scale scenes; consistent evaluation protocols for reconstruction quality, topology, and efficiency across methods are lacking.
- Broader mesh quality metrics: Evaluation uses accuracy and completion within a distance threshold; manifoldness, self-intersection rate, edge length distributions, curvature statistics, and watertightness metrics are not reported.
- Generalization across sensors and settings: The pipeline’s dependence on specific sensor noise models, resolutions, and frame rates is not characterized; guidelines for parameter tuning across sensors/environments are missing.
- Real-time guarantees: Mapping FPS is reported on one sequence and hardware; no analysis of worst-case latency for meshing, update jitter, or QoS under varying scene complexity is provided.
Practical Applications
Immediate Applications
Below are actionable use cases that can be deployed with today’s RGB‑D sensors and existing SLAM stacks, leveraging the paper’s online Gaussian triangulation, plane-based pulling, and freezing strategies.
- Robotics (indoor navigation and manipulation)
- What it enables: Real-time, incrementally updated, watertight triangle meshes for local planning, collision checking, grasp planning, and occlusion-aware perception.
- Potential tools/products/workflows:
- ROS2 node that publishes an incremental triangle mesh (plus normals) alongside point clouds; direct integration with MoveIt, NaviGator, or local planners via mesh/voxelization.
- On-device mapping stack for mobile manipulators and warehouse AGVs that need fast, memory-efficient maps.
- Dependencies/assumptions:
- Requires RGB‑D input or depth from LiDAR; known or reliably estimated poses (from an external SLAM/VIO).
- GPU or strong edge compute for ~real-time optimization; indoor scenes favored.
- AR/VR/MR spatial mapping and occlusion (HMDs, tablets with LiDAR)
- What it enables: High-fidelity, incremental meshes for consistent occlusion, physics, and scene understanding; faster updates and less memory than TSDF-based pipelines.
- Potential tools/products/workflows:
- Unity/Unreal plugin replacing or augmenting ARKit/ARCore/HoloLens meshing; feed explicit meshes directly to occlusion and physics engines.
- On-device apps for interior design and object placement with stable occlusion and crisp geometry.
- Dependencies/assumptions:
- Depth-equipped devices (HoloLens, iPad Pro LiDAR, depth-capable phones/tablets); accurate tracking; sufficient compute.
- AEC/Construction/BIM “as-built vs. as-designed” verification
- What it enables: On-site, near real-time mesh generation to compare against BIM/CAD models, highlight deviations, and perform progress monitoring without offline post-processing.
- Potential tools/products/workflows:
- Tablet-based scanning workflow: capture → incremental meshing → ICP alignment with BIM → deviation heatmap → field report.
- Cloud synchronization of frozen mesh chunks for multi-day projects with limited on-device memory.
- Dependencies/assumptions:
- Indoor environments with adequate texture/depth visibility; pose estimation from SLAM or total station; robust normal estimation for planar constraints.
- Digital twins and facility management
- What it enables: Continuous updates to building/facility meshes with compact, watertight geometry that’s directly streamable to asset systems.
- Potential tools/products/workflows:
- Incremental mesh “patch” publishing service (microservice) that streams updates to a twin platform (e.g., via gRPC/WebSockets).
- Scheduled re-scans that freeze stable regions to control cost and data volume.
- Dependencies/assumptions:
- Persistent sensor coverage; policy-compliant storage/processing of indoor scans; compute budget for on-prem or edge.
- Industrial inspection and QA
- What it enables: Rapid, high-detail mesh capture of machinery/assemblies to detect misalignments, missing parts, or deformation during assembly lines or maintenance.
- Potential tools/products/workflows:
- Workcell-mounted RGB‑D scanners producing incremental meshes; automatic deviation checks against nominal CAD; alerts/stop-the-line triggers.
- Dependencies/assumptions:
- Controlled lighting and non-reflective surfaces; fixed rigs or reliable robot-mounted pose estimation.
- Cultural heritage digitization and museums
- What it enables: On-site scanning that outputs clean, watertight meshes ready for archiving/online viewing without heavy offline processing.
- Potential tools/products/workflows:
- Guided scanning assistant that indicates coverage, freezes stable regions, and exports ready-to-publish meshes.
- Dependencies/assumptions:
- Depth-capable capture; care with reflective/transparent artifacts; possibly multi-scan sessions for occluded areas.
- Telepresence and remote operations
- What it enables: Streaming of incremental meshes for remote users/operators (e.g., site walkthroughs, remote inspection) with less bandwidth than raw point clouds or video for geometry.
- Potential tools/products/workflows:
- WebRTC service that streams mesh deltas from the freezing-based pipeline; clients render with lightweight mesh viewers.
- Dependencies/assumptions:
- Stable network; server-side GPU or client-side fallback for rendering; data privacy compliance.
- 3D content creation and game pipelines
- What it enables: Rapid capture-to-asset workflows producing lightweight, “game-ready” meshes with reduced cleanup versus TSDF or NeRF post-processing.
- Potential tools/products/workflows:
- DCC plugins (Blender/Maya) importing live mesh streams; asset QC tools that exploit normals and mesh topology for auto-simplification.
- Dependencies/assumptions:
- Suitable surface coverage; manual cleanup for specular/transparent materials remains likely.
- Academic research and dataset creation
- What it enables: A baseline for incremental mesh reconstruction with both high-quality rendering and explicit geometry; useful for benchmarking mapping, tracking, and perception.
- Potential tools/products/workflows:
- Open-source reference implementation; dataset curation pipelines that generate per-frame mesh updates and ground truth for downstream tasks.
- Dependencies/assumptions:
- Reproducible capture with RGB‑D sensors; standardized evaluation protocols.
- Consumer/home applications (daily life)
- What it enables: Quick room scanning for furniture placement, energy audits (e.g., measuring gaps), DIY renovation planning with accurate measurements.
- Potential tools/products/workflows:
- Mobile app for room scanning using LiDAR phones/tablets; export watertight meshes to floor-planning or AR furniture apps.
- Dependencies/assumptions:
- Device with depth sensor; reasonable lighting; acceptance of in-home scanning/privacy controls.
Long-Term Applications
These require further R&D, scaling, new sensors, or integration beyond the paper’s current RGB‑D assumption.
- RGB-only online reconstruction (no depth)
- Vision: Extend plane-based pulling and triangulation to monocular or multi-view RGB with learned depth/normal priors.
- Potential tools/products/workflows:
- Smartphone-only AR mapping that outputs high-quality meshes without LiDAR; cloud-assisted triangulation for low-power devices.
- Dependencies/assumptions:
- Robust learned depth/normal estimation; drift-resistant SLAM; handling low-texture/reflective surfaces.
- Outdoor and large-scale mapping (city blocks, infrastructure)
- Vision: Scale Gaussian triangulation and freezing to unbounded scenes with sparse, noisy depth (e.g., LiDAR) and varied lighting.
- Potential tools/products/workflows:
- Mobile mapping rigs for utilities/energy (plants, substations, pipelines); fleet-based asset digitization with mesh patch streaming.
- Dependencies/assumptions:
- GPS/INS-aided SLAM; robust outlier handling; distributed compute/storage; weather and surface variability resilience.
- Dynamic scene reconstruction (time-varying meshes)
- Vision: Handle moving objects and deformable scenes with per-region temporal models; maintain stability via local remeshing.
- Potential tools/products/workflows:
- Service robots operating in human environments (stores/hospitals) that require real-time, up-to-date meshes for planning and safety.
- Dependencies/assumptions:
- Reliable motion segmentation/association; temporal consistency constraints; real-time compute budget.
- Semantic meshes for task-aware robotics and AR
- Vision: Fuse online semantic labels with the triangulated mesh for object-level reasoning, affordances, and scene queries.
- Potential tools/products/workflows:
- Semantic digital twins; AR overlays with object-aware occlusion and interactions; robot pick-and-place with category-level priors.
- Dependencies/assumptions:
- On-the-fly semantic segmentation; robust fusion of semantics with geometry; consistent taxonomy.
- Edge deployment on mobile SoCs and embedded GPUs
- Vision: Highly optimized kernels and memory-aware scheduling to reach real-time on smartphones, AR headsets, and micro-robots.
- Potential tools/products/workflows:
- AR headset firmware modules; consumer vacuum/mower robots with detailed home meshes; drone mapping payloads.
- Dependencies/assumptions:
- Kernel fusion, quantization, and reduced precision; thermal/power constraints; tailored sensor drivers.
- Multi-agent collaborative mapping
- Vision: Mesh patch merging across agents using freezing, conflict resolution, and global consistency enforcement.
- Potential tools/products/workflows:
- Teams of robots/drones incrementally building shared meshes for warehouses or disaster sites.
- Dependencies/assumptions:
- Robust inter-robot localization; mesh diff/merge protocols; network bandwidth and synchronization.
- Healthcare and surgical AR
- Vision: Intraoperative mapping (endoscopy, laparoscopy) with online triangulation for guidance and occlusion; pre/post-operative comparisons.
- Potential tools/products/workflows:
- Surgical navigation systems integrating real-time meshes with pre-op scans; training simulators with high-fidelity geometry.
- Dependencies/assumptions:
- Regulatory approvals; specialized sensors; handling soft tissue dynamics, fluids, and specularity.
- Policy and standards for 3D scanning and data governance
- Vision: Best-practice guidelines for compact, privacy-preserving real-time mesh capture in public and private spaces; interoperability standards for mesh patch streaming.
- Potential tools/products/workflows:
- Open standards for “incremental mesh delta” formats; on-device redaction (e.g., face/logo blurring) applied directly to the mesh.
- Dependencies/assumptions:
- Stakeholder alignment (industry, regulators); privacy risk assessments; standard APIs and codecs.
- Asset valuation and insurance (built environment)
- Vision: Rapid, periodic re-scans producing consistent meshes to assess changes, risks, or damage claims.
- Potential tools/products/workflows:
- Inspections with tablet-based mapping that generates report-ready geometry and measurements.
- Dependencies/assumptions:
- Acceptance of 3D scans as evidence; standardized measurement tolerances; coverage completeness.
- High-fidelity teleoperation and digital workforce
- Vision: Low-latency mesh streaming for remote robot teleop and digital twins that require accurate surface contact reasoning.
- Potential tools/products/workflows:
- Operator UIs with real-time mesh overlays; sim-to-real transfer pipelines using online meshes for environment mirrors.
- Dependencies/assumptions:
- Low-latency comms; robust synchronization with robot state; safety and failover mechanisms.
Notes on Feasibility and Cross-Cutting Assumptions
- Sensor modality: The presented system relies on RGB‑D inputs and known/estimated camera poses; extension to RGB-only is a research item.
- Compute: Real-time performance in the paper uses an RTX 3090; embedded/edge deployment requires further optimization.
- Scene type: Best suited for indoor or controlled environments with good coverage; reflective/transparent surfaces and unobserved regions remain challenging.
- Integration: Most real-world deployments will pair the method with a SLAM/VIO stack for pose estimation and a task-specific consumer (planner, AR engine, BIM tool).
- Parameters and stability: Geometry selection threshold (τ), opacity sparsity, and plane-based pulling require tuning; multi-view coverage is critical for mesh completeness.
Glossary
- Accuracy Ratio: The percentage of reconstructed points within a specified distance threshold from the ground truth; used to quantify geometric precision. Example: "Accuracy Ratio (cm)"
- Adaptive sampling: A strategy that varies sampling density based on scene content to improve efficiency or quality. Example: "models scenes as radiance fields with adaptive sampling."
- Alpha-blending: A rendering technique that composites semi-transparent elements by accumulating colors and opacities along the view ray. Example: "through the standard alpha-blending proposed in \cite{kerbl20233d},"
- Angle-based greedy strategy: A local triangulation procedure that orders neighbors by angle and forms triangles greedily to construct meshes. Example: "We employ an angle-based greedy strategy to reconstruct the triangular meshes:"
- Ball pivoting algorithm: A surface reconstruction method that rolls a virtual ball over points to form triangles. Example: "we employ the ball pivoting algorithm~\cite{bernardini2002ball} to generate meshes from Gaussian balls."
- Compressed octree: A memory-efficient hierarchical spatial index for fast neighbor searches in 3D. Example: "we employ a compressed octree."
- Completion Ratio: The fraction of ground-truth surface points that are within a threshold of the reconstruction; measures completeness. Example: "Completion Ratio (cm)"
- Dense Geometric Gaussian representation: A surface-oriented set of Gaussians tailored for both rendering and direct meshing. Example: "The dense geometric Gaussian representation is progressively refined by loss-guided densification and optimization against depth, color, and geometric constraints."
- Differentiable rendering: Rendering formulations that are amenable to gradient-based optimization of scene parameters. Example: "Geometry-Aware Differentiable Rendering."
- Dual-branch rendering strategy: Joint optimization that uses both color and geometry cues in separate rendering branches. Example: "We adopt a dual-branch rendering strategy to optimize both photometric and geometric consistency."
- Gaussian surfels: Flattened, oriented Gaussian primitives that approximate local surface patches for rendering and meshing. Example: "incremental update from Gaussian surfels;"
- Gaussian Triangulation: Direct mesh construction by triangulating a set of surface-aligned Gaussian primitives. Example: "Gaussian Triangulation"
- Geometric Gaussian Set: The subset of Gaussians deemed reliable for surface reconstruction based on opacity and depth fidelity. Example: "the Geometric Gaussian Set "
- Isotropic mesh optimization: Remeshing that regularizes triangle shapes and sizes to be uniform for improved mesh quality. Example: "Following the isotropic mesh optimization approach proposed by Botsch et al.~\cite{botsch2004remeshing},"
- k-nearest neighbors (k-NN): The set of k closest points to a query point, commonly used to estimate local geometry. Example: "the -nearest neighbors in 3D space"
- Laplacian smoothing: A mesh smoothing technique that relocates vertices based on neighbor averages to reduce noise. Example: "via Laplacian smoothing."
- Marching Cubes: A classic algorithm that extracts triangle meshes from scalar fields on voxel grids. Example: "extracted by Marching Cubes~\cite{lorensen1987marching}"
- Marching Tetrahedra: An alternative to Marching Cubes that operates on tetrahedral decompositions for surface extraction. Example: "followed by surface extraction using Marching Tetrahedra."
- Neural Radiance Field (NeRF): A neural implicit scene representation that models view-dependent appearance and density. Example: "Neural Radiance Field (NeRF)~\cite{mildenhall2021nerf}"
- Normal consistency loss: A penalty that aligns estimated surface normals with reference normals to improve geometric fidelity. Example: "we employ a normal consistency loss defined as"
- Opacity Constraint: A regularization that pushes Gaussian opacities toward solid surfaces for mesh-ready geometry. Example: "Opacity Constraint: To approximate hard physical geometry, we enforce high opacity ()."
- Oriented point cloud: A point cloud where each point is equipped with a surface normal, aiding surface reconstruction. Example: "from a continuous stream of oriented point clouds."
- Plane-based pulling constraint: A geometric alignment term that attracts Gaussian centers to local planes inferred from data. Example: "we enforce a plane-based pulling constraint"
- Planar Elliptical Surfels: Elliptical, near-planar Gaussian elements aligned with local tangent planes to represent surfaces. Example: "Planar Elliptical Surfels:"
- Poisson surface reconstruction: An implicit reconstruction method that solves a Poisson equation from oriented points to produce a watertight mesh. Example: "we adopt Poisson surface reconstruction in place of the proposed direct triangulation."
- Structure-from-Motion (SfM): A pipeline that reconstructs camera poses and sparse 3D points from images. Example: "Structure-from-Motion (SfM)"
- Tangent plane: The local planar approximation of a surface at a point, used to align surfels and project neighbors. Example: "align with the local tangent plane"
- Truncated Signed Distance Field (TSDF): A volumetric representation that stores truncated distances to the nearest surface for fusion-based reconstruction. Example: "Truncated Signed Distance Field (TSDF)"
- Volumetric integration: The incremental fusion of depth observations into a volumetric grid (e.g., TSDF). Example: "volumetric integration of the Truncated Signed Distance Field (TSDF)"
- Volumetric rendering: Rendering by integrating radiance and opacity along rays through a volume. Example: "the volumetric rendering is computationally intensive"
- Voxel resolution: The spatial discretization granularity of a volumetric grid, limiting detail and memory usage. Example: "fixed voxel resolution"
- Watertight meshes: Meshes with no holes or gaps, forming closed surfaces suitable for downstream tasks. Example: "By directly extracting watertight meshes from these primitives"
- Zero level set: The set of points where an implicit function (e.g., SDF) evaluates to zero, representing the surface. Example: "zero level set "







