- The paper introduces SLNet, a hierarchical 3D point cloud backbone combining parameter-free adaptive geometric embedding, lightweight channel modulation, and efficient local aggregation; its smallest model achieves 93.64% accuracy on ModelNet40 with 0.14M parameters.
- SLNet delivers near-baseline performance across classification, few-shot learning, and part segmentation while reducing model size and computation by up to one to two orders of magnitude, including 84.25% accuracy on ScanObjectNN and 85.53% part-segmentation IoU on ShapeNetPart.
- The paper proposes NetScore⁺, which adds latency and memory to accuracy, parameter, and FLOP measurements; the metric correlates with throughput at Spearman ρ > 0.9 and highlights SLNet’s suitability for resource-constrained deployment despite lower scene-segmentation accuracy than large transformers.
SLNet is a lightweight hierarchical backbone for 3D point cloud recognition that combines a fully parameter-free geometric embedding with an ultra-low-cost channel modulation mechanism, targeting deployment on resource-constrained hardware. The paper's central claim is that careful design of geometric encoding, feature modulation, and local aggregation allows a model with as few as 0.14M parameters to remain competitive with substantially larger supervised baselines across classification, few-shot learning, part segmentation, and scene segmentation. The authors also introduce NetScore+, a deployability metric extending NetScore by incorporating latency and peak memory.
Architecture
The backbone is organized into three variants sharing a common four-stage hierarchy built on FPS sampling, kNN grouping (K=32 for ModelNet40/ShapeNetPart; K=24 for ScanObjectNN), parameter-free relative-feature normalization, and shared residual MLPs (Light Residual Blocks with a fixed channel width ratio r=0.25). SLNet-S uses 0.14M parameters and 0.31 GFLOPs; SLNet-M uses 0.55M and 1.22 GFLOPs. For scene segmentation, SLNet-T replaces NAPE with a learned linear projection (6→64 channels) to accommodate XYZ+RGB input, replaces MLP stages with local Point Transformer v1 attention at all four encoder stages, and grows to 2.5M parameters and 6.5 GFLOPs. Classification heads use global adaptive max-pooling followed by a two-layer MLP; part segmentation adds a U-Net decoder with inverse-distance-weighted interpolation and multi-scale GMP fusion; scene segmentation uses a lightweight FP decoder trained with inverse-square-root weighted cross-entropy and label smoothing.
Nonparametric Adaptive Point Embedding
NAPE maps raw XYZ coordinates to a D-dimensional feature without any learnable parameters through three steps. Global dispersion σglobal, the mean per-axis standard deviation, estimates object scale in a translation-invariant manner. The kernel bandwidth adapts to this scale via σadapt=σ0(1+σglobal) with a fixed base bandwidth σ0=0.4. Gaussian RBF and cosine bases are then evaluated on a uniform grid per axis and blended through a sigmoid gate β=sigmoid(γ(σglobal−b)), so that localized Gaussian responses dominate for small objects while smoother cosine responses dominate for larger ones. All hyperparameters are fixed across datasets and model sizes, which strengthens the claim that the encoding generalizes without tuning.
Geometric Modulation Unit
GMU applies a per-channel affine transformation Yb,d,n=αdXb,d,n+βd with only K=320 learnable scalars (32 for SLNet-S, 64 for SLNet-M). The ablation shows GMU placement matters: scale-then-shift (K=321) applied immediately after the embedding recovers 0.5 pp on both variants (93.13→93.64% and 93.44→93.92% on ModelNet40), whereas post-sampling/grouping placement, reversed order, or double application consistently underperforms. Notably, GMU is omitted entirely for ScanObjectNN based on ablation results—an honest concession that the module is not universally beneficial.
NetScoreK=322
The paper extends NetScore (Wong, 2018) with a latency–memory term:
K=323
where K=324 is accuracy in percent, K=325 parameters, K=326 FLOPs, K=327 inference time, and K=328 memory footprint. Across 12 architectures, NetScoreK=329 correlates strongly with measured throughput (Spearman K=240), supporting its use as a deployment-oriented composite metric. Evaluations span RTX 3090 and Jetson Orin Nano platforms, covering 12 metric–dataset–hardware axes in which SLNet-S occupies the outermost (best) position throughout.
Classification results
On ModelNet40, SLNet-S achieves 93.64% OA—surpassing PointMLP-elite (93.28%) with K=241 fewer parameters—and posts the highest NetScore (92.42) and NetScoreK=242 (87.71) among all evaluated methods. SLNet-M reaches 93.92%, exceeding PointMLP's reproduced result (93.66%) with K=243 fewer parameters and K=244 fewer FLOPs. On ModelNet-R, a re-annotated benchmark with corrected labels, SLNet-S/M reach 94.53%/94.81%. On ScanObjectNN PB-T50-RS, SLNet-M attains 84.25% OA within 1.15 pp of PointMLP (85.40%) using K=245 fewer parameters and K=246 fewer FLOPs. These results substantiate the paper's core efficiency claim: near-parity accuracy at one to two orders of magnitude smaller parameter budgets.
Segmentation results
On ShapeNetPart, SLNet-S achieves 85.21% ins-IoU with the highest NetScoreK=247 of all evaluated methods (66.81/60.31); SLNet-M reaches 85.53% ins-IoU at 3.72 ms per cloud (2048 points), K=248 faster than CurveNet at 1.1 pp lower IoU. On S3DIS Area 5, SLNet-T achieves 58.2% mIoU, 85.3% OA, and 65.4% mAcc with 2.5M parameters—roughly K=249 smaller than Point Transformer V3. The paper is explicit that absolute mIoU trails heavier transformer baselines (PT: 70.4%; ConDAF: 73.5%), but argues SLNet-T attains the highest NetScore (58.5 vs. 57.5 for PT), reflecting its design goal of prioritizing efficiency under a strict parameter budget rather than maximizing raw accuracy.
In few-shot classification on ModelNet40, SLNet-M reaches 95.0% (5-way 20-shot) and 94.0% (10-way 20-shot) without large-scale pretraining. In the 10-way 20-shot setting, both parametric variants surpass all non-parametric baselines including NPNet (87.6%), improving by roughly 6 pp—a somewhat counterintuitive result given that non-parametric methods typically dominate low-data regimes.
Ablation findings
Several ablations carry independent weight. Embedding choice: NAPE alone outperforms learned MLP, pure Gaussian, and pure cosine embeddings by 0.9–1.1 pp, indicating adaptive basis blending is more effective than any single basis. Neighborhood size: r=0.250 is optimal for clean ModelNet40 data while r=0.251 suits cluttered ScanObjectNN; performance degrades noticeably at r=0.252 (−0.75 pp), consistent with noisy data introducing irrelevant context. Residual bottleneck: r=0.253 is Pareto-optimal—the r=0.254 setting loses 1.6 pp while saving only 0.03M parameters, and wider bottlenecks add cost without gain. Sampling: FPS beats random sampling by 0.3 pp on ModelNet40 and 0.6 pp on ScanObjectNN, the larger gap attributed to FPS maintaining spatial coverage under occlusion. Scene-level design: replacing MLP aggregation with full local attention yields a +9.6 pp mIoU gain (43.68→53.24%), confirming shared MLPs lack relational capacity to disambiguate adjacent categories such as wall, column, and beam; inverse-square-root weighted CE adds +2.7 pp over plain CE and outperforms focal loss; EMA is the most impactful regularizer (−0.9 pp when removed). A half-width configuration reduces parameters and FLOPs by 75% (to 0.62M / 1.64G) at a cost of 5.5 pp mIoU, identified as an attractive operating point for extremely constrained deployments. Qualitative gradient saliency shows NAPE producing sharper, semantically focused responses than DGCNN EdgeConv features despite operating at the input layer.
Limitations and open questions
The paper concedes several boundaries explicitly. Absolute accuracy on S3DIS remains well below state-of-the-art transformer baselines, so SLNet-T's advantage holds only under the NetScore framing rather than raw segmentation quality. GMU's benefit is dataset-dependent (omitted for ScanObjectNN), and neighborhood size must be tuned per dataset, weakening any claim of a single universally optimal configuration. The OneCycle-plus-slow-EMA combination diverged during training, indicating sensitivity of the training recipe to schedule interactions. Few-shot gains are reported only for specific shot configurations, and the strong correlation claimed for NetScorer=0.255 rests on 12 architectures evaluated under particular batch sizes and input resolutions; whether the metric generalizes to other regimes is left unexamined. The paper also does not evaluate robustness to rotation or density variation beyond what ScanObjectNN's augmentation implicitly covers.
Conclusion
SLNet demonstrates that a backbone combining parameter-free adaptive geometric embedding (NAPE), minimal affine modulation (GMU), and efficient local aggregation can match or approach much larger baselines on classification, part segmentation, and few-shot tasks while reducing parameter counts by up to two orders of magnitude. Its extension to scene segmentation trades absolute mIoU for a favorable efficiency profile quantified through the proposed NetScorer=0.256 metric. The main open question the work leaves is whether the NAPE/GMU design principles transfer to outdoor LiDAR and higher-density inputs, where object-scale statistics and clutter characteristics differ substantially from the indoor benchmarks evaluated here.