Papers
Topics
Authors
Recent
Search
2000 character limit reached

SLNet: A Super-Lightweight Geometry-Adaptive Network for 3D Point Cloud Recognition

Published 8 Mar 2026 in cs.CV, cs.LG, and cs.RO | (2603.07454v1)

Abstract: We present SLNet, a lightweight backbone for 3D point cloud recognition designed to achieve strong performance without the computational cost of many recent attention, graph, and deep MLP based models. The model is built on two simple ideas: NAPE (Nonparametric Adaptive Point Embedding), which captures spatial structure using a combination of Gaussian RBF and cosine bases with input adaptive bandwidth and blending, and GMU (Geometric Modulation Unit), a per channel affine modulator that adds only 2D learnable parameters. These components are used within a four stage hierarchical encoder with FPS+kNN grouping, nonparametric normalization, and shared residual MLPs. In experiments, SLNet shows that a very small model can still remain highly competitive across several 3D recognition tasks. On ModelNet40, SLNet-S with 0.14M parameters and 0.31 GFLOPs achieves 93.64% overall accuracy, outperforming PointMLP-elite with 5x fewer parameters, while SLNet-M with 0.55M parameters and 1.22 GFLOPs reaches 93.92%, exceeding PointMLP with 24x fewer parameters. On ScanObjectNN, SLNet-M achieves 84.25% overall accuracy within 1.2 percentage points of PointMLP while using 28x fewer parameters. For large scale scene segmentation, SLNet-T extends the backbone with local Point Transformer attention and reaches 58.2% mIoU on S3DIS Area 5 with only 2.5M parameters, more than 17x fewer than Point Transformer V3. We also introduce NetScore+, which extends NetScore by incorporating latency and peak memory so that efficiency can be evaluated in a more deployment oriented way. Across multiple benchmarks and hardware settings, SLNet delivers a strong overall balance between accuracy and efficiency. Code is available at: https://github.com/m-saeid/SLNet.

Summary

  • The paper introduces SLNet, a hierarchical 3D point cloud backbone combining parameter-free adaptive geometric embedding, lightweight channel modulation, and efficient local aggregation; its smallest model achieves 93.64% accuracy on ModelNet40 with 0.14M parameters.
  • SLNet delivers near-baseline performance across classification, few-shot learning, and part segmentation while reducing model size and computation by up to one to two orders of magnitude, including 84.25% accuracy on ScanObjectNN and 85.53% part-segmentation IoU on ShapeNetPart.
  • The paper proposes NetScore⁺, which adds latency and memory to accuracy, parameter, and FLOP measurements; the metric correlates with throughput at Spearman ρ > 0.9 and highlights SLNet’s suitability for resource-constrained deployment despite lower scene-segmentation accuracy than large transformers.

SLNet is a lightweight hierarchical backbone for 3D point cloud recognition that combines a fully parameter-free geometric embedding with an ultra-low-cost channel modulation mechanism, targeting deployment on resource-constrained hardware. The paper's central claim is that careful design of geometric encoding, feature modulation, and local aggregation allows a model with as few as 0.14M parameters to remain competitive with substantially larger supervised baselines across classification, few-shot learning, part segmentation, and scene segmentation. The authors also introduce NetScore+^{+}, a deployability metric extending NetScore by incorporating latency and peak memory.

Architecture

The backbone is organized into three variants sharing a common four-stage hierarchy built on FPS sampling, kNN grouping (K=32K{=}32 for ModelNet40/ShapeNetPart; K=24K{=}24 for ScanObjectNN), parameter-free relative-feature normalization, and shared residual MLPs (Light Residual Blocks with a fixed channel width ratio r=0.25r{=}0.25). SLNet-S uses 0.14M parameters and 0.31 GFLOPs; SLNet-M uses 0.55M and 1.22 GFLOPs. For scene segmentation, SLNet-T replaces NAPE with a learned linear projection (6→64 channels) to accommodate XYZ+RGB input, replaces MLP stages with local Point Transformer v1 attention at all four encoder stages, and grows to 2.5M parameters and 6.5 GFLOPs. Classification heads use global adaptive max-pooling followed by a two-layer MLP; part segmentation adds a U-Net decoder with inverse-distance-weighted interpolation and multi-scale GMP fusion; scene segmentation uses a lightweight FP decoder trained with inverse-square-root weighted cross-entropy and label smoothing.

Nonparametric Adaptive Point Embedding

NAPE maps raw XYZ coordinates to a DD-dimensional feature without any learnable parameters through three steps. Global dispersion σglobal\sigma_{\mathrm{global}}, the mean per-axis standard deviation, estimates object scale in a translation-invariant manner. The kernel bandwidth adapts to this scale via σadapt=σ0(1+σglobal)\sigma_{\mathrm{adapt}} = \sigma_0(1 + \sigma_{\mathrm{global}}) with a fixed base bandwidth σ0=0.4\sigma_0 = 0.4. Gaussian RBF and cosine bases are then evaluated on a uniform grid per axis and blended through a sigmoid gate β=sigmoid(γ(σglobal−b))\beta = \mathrm{sigmoid}(\gamma(\sigma_{\mathrm{global}} - b)), so that localized Gaussian responses dominate for small objects while smoother cosine responses dominate for larger ones. All hyperparameters are fixed across datasets and model sizes, which strengthens the claim that the encoding generalizes without tuning.

Geometric Modulation Unit

GMU applies a per-channel affine transformation Yb,d,n=αdXb,d,n+βd\mathbf{Y}_{b,d,n} = \alpha_d \mathbf{X}_{b,d,n} + \beta_d with only K=32K{=}320 learnable scalars (32 for SLNet-S, 64 for SLNet-M). The ablation shows GMU placement matters: scale-then-shift (K=32K{=}321) applied immediately after the embedding recovers 0.5 pp on both variants (93.13→93.64% and 93.44→93.92% on ModelNet40), whereas post-sampling/grouping placement, reversed order, or double application consistently underperforms. Notably, GMU is omitted entirely for ScanObjectNN based on ablation results—an honest concession that the module is not universally beneficial.

NetScoreK=32K{=}322

The paper extends NetScore (Wong, 2018) with a latency–memory term:

K=32K{=}323

where K=32K{=}324 is accuracy in percent, K=32K{=}325 parameters, K=32K{=}326 FLOPs, K=32K{=}327 inference time, and K=32K{=}328 memory footprint. Across 12 architectures, NetScoreK=32K{=}329 correlates strongly with measured throughput (Spearman K=24K{=}240), supporting its use as a deployment-oriented composite metric. Evaluations span RTX 3090 and Jetson Orin Nano platforms, covering 12 metric–dataset–hardware axes in which SLNet-S occupies the outermost (best) position throughout.

Classification results

On ModelNet40, SLNet-S achieves 93.64% OA—surpassing PointMLP-elite (93.28%) with K=24K{=}241 fewer parameters—and posts the highest NetScore (92.42) and NetScoreK=24K{=}242 (87.71) among all evaluated methods. SLNet-M reaches 93.92%, exceeding PointMLP's reproduced result (93.66%) with K=24K{=}243 fewer parameters and K=24K{=}244 fewer FLOPs. On ModelNet-R, a re-annotated benchmark with corrected labels, SLNet-S/M reach 94.53%/94.81%. On ScanObjectNN PB-T50-RS, SLNet-M attains 84.25% OA within 1.15 pp of PointMLP (85.40%) using K=24K{=}245 fewer parameters and K=24K{=}246 fewer FLOPs. These results substantiate the paper's core efficiency claim: near-parity accuracy at one to two orders of magnitude smaller parameter budgets.

Segmentation results

On ShapeNetPart, SLNet-S achieves 85.21% ins-IoU with the highest NetScoreK=24K{=}247 of all evaluated methods (66.81/60.31); SLNet-M reaches 85.53% ins-IoU at 3.72 ms per cloud (2048 points), K=24K{=}248 faster than CurveNet at 1.1 pp lower IoU. On S3DIS Area 5, SLNet-T achieves 58.2% mIoU, 85.3% OA, and 65.4% mAcc with 2.5M parameters—roughly K=24K{=}249 smaller than Point Transformer V3. The paper is explicit that absolute mIoU trails heavier transformer baselines (PT: 70.4%; ConDAF: 73.5%), but argues SLNet-T attains the highest NetScore (58.5 vs. 57.5 for PT), reflecting its design goal of prioritizing efficiency under a strict parameter budget rather than maximizing raw accuracy.

In few-shot classification on ModelNet40, SLNet-M reaches 95.0% (5-way 20-shot) and 94.0% (10-way 20-shot) without large-scale pretraining. In the 10-way 20-shot setting, both parametric variants surpass all non-parametric baselines including NPNet (87.6%), improving by roughly 6 pp—a somewhat counterintuitive result given that non-parametric methods typically dominate low-data regimes.

Ablation findings

Several ablations carry independent weight. Embedding choice: NAPE alone outperforms learned MLP, pure Gaussian, and pure cosine embeddings by 0.9–1.1 pp, indicating adaptive basis blending is more effective than any single basis. Neighborhood size: r=0.25r{=}0.250 is optimal for clean ModelNet40 data while r=0.25r{=}0.251 suits cluttered ScanObjectNN; performance degrades noticeably at r=0.25r{=}0.252 (−0.75 pp), consistent with noisy data introducing irrelevant context. Residual bottleneck: r=0.25r{=}0.253 is Pareto-optimal—the r=0.25r{=}0.254 setting loses 1.6 pp while saving only 0.03M parameters, and wider bottlenecks add cost without gain. Sampling: FPS beats random sampling by 0.3 pp on ModelNet40 and 0.6 pp on ScanObjectNN, the larger gap attributed to FPS maintaining spatial coverage under occlusion. Scene-level design: replacing MLP aggregation with full local attention yields a +9.6 pp mIoU gain (43.68→53.24%), confirming shared MLPs lack relational capacity to disambiguate adjacent categories such as wall, column, and beam; inverse-square-root weighted CE adds +2.7 pp over plain CE and outperforms focal loss; EMA is the most impactful regularizer (−0.9 pp when removed). A half-width configuration reduces parameters and FLOPs by 75% (to 0.62M / 1.64G) at a cost of 5.5 pp mIoU, identified as an attractive operating point for extremely constrained deployments. Qualitative gradient saliency shows NAPE producing sharper, semantically focused responses than DGCNN EdgeConv features despite operating at the input layer.

Limitations and open questions

The paper concedes several boundaries explicitly. Absolute accuracy on S3DIS remains well below state-of-the-art transformer baselines, so SLNet-T's advantage holds only under the NetScore framing rather than raw segmentation quality. GMU's benefit is dataset-dependent (omitted for ScanObjectNN), and neighborhood size must be tuned per dataset, weakening any claim of a single universally optimal configuration. The OneCycle-plus-slow-EMA combination diverged during training, indicating sensitivity of the training recipe to schedule interactions. Few-shot gains are reported only for specific shot configurations, and the strong correlation claimed for NetScorer=0.25r{=}0.255 rests on 12 architectures evaluated under particular batch sizes and input resolutions; whether the metric generalizes to other regimes is left unexamined. The paper also does not evaluate robustness to rotation or density variation beyond what ScanObjectNN's augmentation implicitly covers.

Conclusion

SLNet demonstrates that a backbone combining parameter-free adaptive geometric embedding (NAPE), minimal affine modulation (GMU), and efficient local aggregation can match or approach much larger baselines on classification, part segmentation, and few-shot tasks while reducing parameter counts by up to two orders of magnitude. Its extension to scene segmentation trades absolute mIoU for a favorable efficiency profile quantified through the proposed NetScorer=0.25r{=}0.256 metric. The main open question the work leaves is whether the NAPE/GMU design principles transfer to outdoor LiDAR and higher-density inputs, where object-scale statistics and clutter characteristics differ substantially from the indoor benchmarks evaluated here.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.