Papers
Topics
Authors
Recent
Search
2000 character limit reached

Point-to-Ellipsoid (POLI) Framework

Updated 12 July 2026
  • POLI is a learned representation that models local point-cloud geometry as per-point Gaussians whose covariance matrices form ellipsoids, capturing uncertainties and anisotropic structure.
  • The framework uses a statistical manifold formulation and a PointNet++-inspired network to predict covariances, which improve key tasks like localization, mapping, and object pose estimation.
  • Experimental results demonstrate that POLI enhances registration success and reduces error rates in LiDAR odometry and scan matching, especially under sparse and low-resolution conditions.

Point-to-Ellipsoid (POLI) denotes a representation and estimation framework in which local point-cloud geometry is modeled by per-point Gaussian distributions whose covariance matrices define ellipsoids. In the formulation introduced in “Learning Point Cloud Geometry as a Statistical Manifold: Theory and Practice” (Lee et al., 11 May 2026), POLI is a deep neural estimator that predicts per-point Gaussian geometry from LiDAR observations in a self-supervised manner. Each point is associated with a covariance matrix in S+3\mathbb{S}_+^3, and the resulting ellipsoid encodes local anisotropy, tangent-plane structure, and uncertainty. The representation is designed for robotic perception tasks such as localization, mapping, object pose estimation, and scan matching, and is intended as a drop-in geometric module rather than as a task-specific reconstruction system (Lee et al., 11 May 2026).

1. Problem setting and conceptual motivation

POLI is motivated by a specific difficulty in LiDAR perception: the observed point cloud is only a sparse, non-uniform, view-dependent sampling of an underlying continuous $3$D surface. Local neighborhoods are therefore shaped not only by the surface itself, but also by sensor angular resolution, distance-dependent sparsity, occlusion, viewpoint, scanning pattern, and partial observability. The central claim is that this makes local geometry estimation difficult when geometry is inferred directly from raw neighborhood statistics (Lee et al., 11 May 2026).

Within this framing, POLI is positioned against two dominant families of prior methods. The first consists of hand-crafted local statistics such as PCA-based covariance estimation, point-to-plane ICP, GICP, PFH/FPFH, and SHOT. These methods estimate geometry directly from observed neighborhoods, and the critique is that they conflate observation statistics with underlying geometry. The second family consists of supervised learning approaches for normal estimation, completion, or densification. These can improve scalability, but often require large labeled datasets, rely on dense scans or accurate labels, delegate geometry reasoning implicitly to the network, lack explicit geometric inductive bias, or produce representations that are difficult to integrate consistently into downstream robotics pipelines (Lee et al., 11 May 2026).

The conceptual response is to represent geometry as a statistical manifold Mg\mathcal{M}_{\mathbf g} induced by a family of Gaussian distributions. Each point is assigned a Gaussian that approximates local surface structure. This replaces a purely descriptive neighborhood statistic with an explicit probabilistic local model, and it makes covariance, anisotropy, and tangent structure first-class geometric quantities rather than by-products of preprocessing (Lee et al., 11 May 2026).

2. Statistical-manifold formulation

The mathematical object underlying POLI is a family of local Gaussians

N(xˉgi,Cgi),\mathcal{N}(\bar{\mathbf{x}}_{\mathbf{g}_i}, \mathbf{C}_{\mathbf{g}_i}),

where xˉgi∈R3\bar{\mathbf{x}}_{\mathbf{g}_i}\in\mathbb{R}^3 is the mean and Cgi∈S+3\mathbf{C}_{\mathbf{g}_i}\in\mathbb{S}_+^3 is the covariance associated with the ii-th local geometry element. The paper states that each Gaussian “defines the probability density that an observed point xi\mathbf{x}_i is measured with respect to a local tangent plane at xˉgi\bar{\mathbf{x}}_{\mathbf{g}_i}” (Lee et al., 11 May 2026).

Under a sensor pose Tp=(Rp,tp)∈SE(3)\mathbf{T_p}=(\mathbf{R_p},\mathbf{t_p})\in SE(3), the same world geometry is transported into the sensor frame by

$3$0

and analogously for a second frame $3$1. This gives the representation the correct rigid-transport behavior across viewpoints (Lee et al., 11 May 2026).

Given two scans

$3$2

with known correspondences and relative transform

$3$3

POLI defines the residual

$3$4

The derivation in the paper yields the Gaussian residual model

$3$5

which is the basis for the point-to-ellipsoid likelihood. The residual is therefore not evaluated with a Euclidean norm but with an anisotropic covariance-shaped Mahalanobis term (Lee et al., 11 May 2026).

The conditional likelihood for all correspondences is written as

$3$6

Because training uses a pose measurement $3$7 rather than exact relative pose, the method introduces a pose likelihood

$3$8

with

$3$9

The resulting energy is

Mg\mathcal{M}_{\mathbf g}0

A Laplace approximation around the mode pose Mg\mathcal{M}_{\mathbf g}1 converts the marginal likelihood into an approximate maximum-likelihood objective over the covariance field (Lee et al., 11 May 2026).

3. Network architecture and learning objective

POLI uses a network inspired by PointNet++. The target point cloud Mg\mathcal{M}_{\mathbf g}2 is processed by Farthest Point Sampling, ball query neighborhoods, centroid-frame normalization, a shared MLP, max pooling, and feature propagation back to original points using distance-weighted interpolation and skip connections. The architecture is explicitly hierarchical and local-feature aware rather than a global latent scene model (Lee et al., 11 May 2026).

For each point Mg\mathcal{M}_{\mathbf g}3, the network predicts six parameters defining a lower-triangular matrix Mg\mathcal{M}_{\mathbf g}4, and the covariance is reconstructed as

Mg\mathcal{M}_{\mathbf g}5

This guarantees Mg\mathcal{M}_{\mathbf g}6 and gives a minimal six-parameter description of a symmetric Mg\mathcal{M}_{\mathbf g}7 covariance matrix (Lee et al., 11 May 2026).

The method is self-supervised because it does not require ground-truth local geometry labels. The supervision signal is induced by the probabilistic model itself: the predicted covariances are optimized according to how well they explain correspondence residuals and pose-consistent alignment. In the notation of the paper, the training loss is

Mg\mathcal{M}_{\mathbf g}8

where Mg\mathcal{M}_{\mathbf g}9 is the Hessian of N(xˉgi,Cgi),\mathcal{N}(\bar{\mathbf{x}}_{\mathbf{g}_i}, \mathbf{C}_{\mathbf{g}_i}),0 at the mode pose (Lee et al., 11 May 2026).

In practice, the inner pose problem is relaxed to

N(xˉgi,Cgi),\mathcal{N}(\bar{\mathbf{x}}_{\mathbf{g}_i}, \mathbf{C}_{\mathbf{g}_i}),1

and this weighted least-squares problem is solved by SDPRLayer. The supplement formulates the relaxed optimization as a homogeneous QCQP with a Shor semidefinite relaxation, enabling a differentiable optimizer that is described as certifiably globally optimal (Lee et al., 11 May 2026).

The paper distinguishes two likelihood formulations. T-Q-C assumes exact transformation and optimizes only the correspondence likelihood. T-T-Q-C incorporates uncertain pose through the pose-likelihood term. The reported comparison states that when exact pose is available, both are similar, but when pose labels are noisy, T-T-Q-C consistently performs better (Lee et al., 11 May 2026).

4. Geometric meaning of the ellipsoid representation

The ellipsoid associated with a point is the confidence region of a local Gaussian. If

N(xˉgi,Cgi),\mathcal{N}(\bar{\mathbf{x}}_{\mathbf{g}_i}, \mathbf{C}_{\mathbf{g}_i}),2

then the eigenvectors provide principal directions of local geometry and the eigenvalues determine the ellipsoid axis lengths. The paper explicitly uses the eigenvector corresponding to the smallest eigenvalue as the surface normal in object pose estimation (Lee et al., 11 May 2026).

This construction gives POLI more structure than a normal-only estimator. A normal supplies one direction; POLI supplies a full covariance, encoding normal direction, tangent-plane extent, anisotropy, and local confidence structure. The paper further states that the covariance captures mainly first-fundamental-form aspects of local geometry rather than curvature, so the ellipsoid should be interpreted as a local probabilistic tangent-structure approximation rather than as a full higher-order surface model (Lee et al., 11 May 2026).

At inference time, the predicted ellipsoids can be used in several ways. They can serve directly as local covariances in GICP and scan matching; they can be eigendecomposed to obtain normals for FPFH-based object pose estimation; and they can be used as Gaussian sampling distributions for scan augmentation. The paper emphasizes that the same output representation integrates into existing pipelines without architectural redesign, and it reports integration into GICP, KISS-ICP, KISS-Matcher, TEASER, FAST-LIO2, and GLIM (Lee et al., 11 May 2026).

A recurrent point in the paper is that the ellipsoid is not merely a noise model on the observed point. It is presented as a probabilistic approximation of local surface geometry. This suggests that the essential contribution of POLI is not simply anisotropic weighting, but the replacement of raw neighborhood statistics by a learned local geometric object (Lee et al., 11 May 2026).

5. Experimental evidence and downstream behavior

The empirical program covers object pose estimation, LiDAR odometry and scan matching, scan augmentation, and plug-and-play integration into existing robotics systems. Across these tasks, the paper reports that POLI improves performance most clearly in sparse and low-resolution regimes (Lee et al., 11 May 2026).

For object pose estimation, the evaluation uses the Stanford 3D Scanning Repository, training on Bunny, Armadillo, Asian Dragon, Lucy, and Thai Statue and testing on Dragon and Happy Buddha. POLI normals are extracted from the smallest-eigenvalue eigenvector and used in an Open3D FPFH pipeline with ROBIN for inlier selection. A registration is counted as successful when rotation error is N(xˉgi,Cgi),\mathcal{N}(\bar{\mathbf{x}}_{\mathbf{g}_i}, \mathbf{C}_{\mathbf{g}_i}),3, and the paper reports that POLI+FPFH achieves higher registration success rates across most settings and sparsity levels, including generalization to unseen test objects (Lee et al., 11 May 2026).

For LiDAR odometry and scan matching on HeLiPR, with training on Roundabout, Town, Riverside, and Bridge and testing on KAIST and DCC, the reported result is that POLI consistently outperformed classical ICP-based approaches across all sequences and achieved performance comparable to dedicated LiDAR odometry systems. The baseline list includes point-to-point ICP, point-to-plane ICP, GICP, VGICP, A-LOAM, F-LOAM, KISS-ICP, and GenZ-ICP (Lee et al., 11 May 2026).

For scan augmentation, the paper gives several explicit improvements. On HeLiPR, KISS-ICP with POLI augmentation changes Bridge from N(xˉgi,Cgi),\mathcal{N}(\bar{\mathbf{x}}_{\mathbf{g}_i}, \mathbf{C}_{\mathbf{g}_i}),4 to N(xˉgi,Cgi),\mathcal{N}(\bar{\mathbf{x}}_{\mathbf{g}_i}, \mathbf{C}_{\mathbf{g}_i}),5, Riverside from N(xˉgi,Cgi),\mathcal{N}(\bar{\mathbf{x}}_{\mathbf{g}_i}, \mathbf{C}_{\mathbf{g}_i}),6 to N(xˉgi,Cgi),\mathcal{N}(\bar{\mathbf{x}}_{\mathbf{g}_i}, \mathbf{C}_{\mathbf{g}_i}),7, Roundabout from N(xˉgi,Cgi),\mathcal{N}(\bar{\mathbf{x}}_{\mathbf{g}_i}, \mathbf{C}_{\mathbf{g}_i}),8 to N(xˉgi,Cgi),\mathcal{N}(\bar{\mathbf{x}}_{\mathbf{g}_i}, \mathbf{C}_{\mathbf{g}_i}),9, Town from xˉgi∈R3\bar{\mathbf{x}}_{\mathbf{g}_i}\in\mathbb{R}^30 to xˉgi∈R3\bar{\mathbf{x}}_{\mathbf{g}_i}\in\mathbb{R}^31, DCC from xˉgi∈R3\bar{\mathbf{x}}_{\mathbf{g}_i}\in\mathbb{R}^32 to xˉgi∈R3\bar{\mathbf{x}}_{\mathbf{g}_i}\in\mathbb{R}^33, and KAIST from xˉgi∈R3\bar{\mathbf{x}}_{\mathbf{g}_i}\in\mathbb{R}^34 to xˉgi∈R3\bar{\mathbf{x}}_{\mathbf{g}_i}\in\mathbb{R}^35. For global registration, KISS-Matcher on HeLiPR changes from xˉgi∈R3\bar{\mathbf{x}}_{\mathbf{g}_i}\in\mathbb{R}^36 raw to xˉgi∈R3\bar{\mathbf{x}}_{\mathbf{g}_i}\in\mathbb{R}^37 with POLI, and TEASER changes from xˉgi∈R3\bar{\mathbf{x}}_{\mathbf{g}_i}\in\mathbb{R}^38 raw to xˉgi∈R3\bar{\mathbf{x}}_{\mathbf{g}_i}\in\mathbb{R}^39 with POLI. On KITTI, KISS-Matcher changes from Cgi∈S+3\mathbf{C}_{\mathbf{g}_i}\in\mathbb{S}_+^30 raw to Cgi∈S+3\mathbf{C}_{\mathbf{g}_i}\in\mathbb{S}_+^31 with POLI, and TEASER changes from Cgi∈S+3\mathbf{C}_{\mathbf{g}_i}\in\mathbb{S}_+^32 raw to Cgi∈S+3\mathbf{C}_{\mathbf{g}_i}\in\mathbb{S}_+^33 with POLI (Lee et al., 11 May 2026).

The strongest gains are reported under reduced LiDAR resolution. For TEASER on KITTI, POLI changes the Cgi∈S+3\mathbf{C}_{\mathbf{g}_i}\in\mathbb{S}_+^34-beam case from Cgi∈S+3\mathbf{C}_{\mathbf{g}_i}\in\mathbb{S}_+^35 to Cgi∈S+3\mathbf{C}_{\mathbf{g}_i}\in\mathbb{S}_+^36 and the Cgi∈S+3\mathbf{C}_{\mathbf{g}_i}\in\mathbb{S}_+^37-beam case from Cgi∈S+3\mathbf{C}_{\mathbf{g}_i}\in\mathbb{S}_+^38 to Cgi∈S+3\mathbf{C}_{\mathbf{g}_i}\in\mathbb{S}_+^39, while the ii0-beam case shows a slight decrease from ii1 to ii2. This is consistent with the stated motivation that learned geometry should be most useful when local neighborhoods are poor direct estimators of surface structure (Lee et al., 11 May 2026).

The plug-and-play evaluations also report improvements. On DCC, FAST-LIO2 changes from RPE trans ii3 to ii4, RPE rot ii5 to ii6, APE trans ii7 to ii8, and APE rot ii9 to xi\mathbf{x}_i0. On the same sequence, GLIM changes from RPE trans xi\mathbf{x}_i1 to xi\mathbf{x}_i2, RPE rot xi\mathbf{x}_i3 to xi\mathbf{x}_i4, APE trans xi\mathbf{x}_i5 to xi\mathbf{x}_i6, and APE rot xi\mathbf{x}_i7 to xi\mathbf{x}_i8 (Lee et al., 11 May 2026).

The reported runtime is 73.9 ms per training step for about 16,000 input points, and inference is 10 ms per 10,000-point scan on an NVIDIA RTX 4060 Ti GPU. These figures are part of the argument that the representation is not only mathematically explicit but also operationally practical (Lee et al., 11 May 2026).

6. Scope, limitations, and relation to adjacent problem families

The POLI acronym refers specifically to the learned covariance estimator of (Lee et al., 11 May 2026), but the phrase “point-to-ellipsoid” appears across several adjacent literatures. Exact centered ellipsoid interpolation studies whether there exists an origin-symmetric ellipsoid passing through points xi\mathbf{x}_i9, equivalently whether there exists xˉgi\bar{\mathbf{x}}_{\mathbf{g}_i}0 such that xˉgi\bar{\mathbf{x}}_{\mathbf{g}_i}1 for all xˉgi\bar{\mathbf{x}}_{\mathbf{g}_i}2 (Tulsiani et al., 2023). Polyellipsoid geometry studies weighted sums of distances to multiple foci and gives containment through conditions of the form

xˉgi\bar{\mathbf{x}}_{\mathbf{g}_i}3

which is a point-membership test rather than a learned local geometric estimator (Blanco et al., 2019). Ellipsoid pose estimation from one ellipse–ellipsoid correspondence uses the tangent cone from an external point to an ellipsoid, with cone matrix

xˉgi\bar{\mathbf{x}}_{\mathbf{g}_i}4

and addresses silhouette and pose constraints rather than local LiDAR geometry learning (Gaudillière et al., 2022). Ellipsoid-based LiDAR registration in EllipseLIO uses per-point ellipsoids derived from tensor voting, but the residual is an ellipsoid-conditioned blend of point-to-line, point-to-plane, and point-to-point targets rather than a literal point-to-ellipsoid-surface distance (Border et al., 20 May 2026).

This broader comparison clarifies a common misconception. POLI is not a paper on exact shortest distance from a query point to the boundary of a fixed ellipsoid, and it is not a nearest-point projection method. It is also not a full surface-reconstruction method and not a curvature estimator. The limitations stated in the paper are that OOD generalization is imperfect, only low-order local geometry is captured, and an exact differentiable solution to the full pose objective remains challenging (Lee et al., 11 May 2026).

At the same time, the broader literature shows that point-to-ellipsoid geometry includes containment, interpolation, tangency, visibility, and registration. This suggests that POLI occupies one specific position within a larger technical family: it turns local surface reasoning into estimation of per-point Gaussian ellipsoids, while adjacent works address exact quadratic feasibility (Tulsiani et al., 2023), multifocal containment (Blanco et al., 2019), tangent-cone projection geometry (Gaudillière et al., 2022), or ellipsoid-conditioned registration objectives (Border et al., 20 May 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Point-to-Ellipsoid (POLI).