- The paper introduces a GPU-based processing pipeline for 3D convex hull filtering, leveraging CUDA, Tensor, and Ray Tracing cores. The method achieves up to $210 imes$ speed and $75 imes$ energy efficiency improvements over CPU for uniform point distributions; performance graphifies to floor for the worst case scenario (sphere).
- The new approach relies on ray tracing for efficient containment checks and TC operations on prefix scans. The process involves several computational phases executed primarily on the GPU with exception for the final hull computation that remains on the CPU, though this is highlighted as a potential area for future optimization.
- The workload of the RT cores greatly varies by the distribution shape of the input points. Uniform distributions are favorable for GPU-based methods. Real-world data potentially provides a challenge because this method is not well tested for this. The paper suggests that further development may push the cut-off curve closer to ideal with optimized GPU dynamic BHI budgeting strategies.
Overview
This paper presents a GPU-based preprocessing filter for the 3D convex hull problem that combines three heterogeneous compute resources available on modern NVIDIA GPUs: CUDA cores, Tensor Cores (TC), and Ray Tracing (RT) cores. The filter constructs a 24-face filtering polyhedron from 14 extreme points of the input set, discards interior points via hardware-accelerated ray/polygon intersection, and compacts the surviving candidate set using a TC-accelerated prefix scan. Evaluated against the Pseudohull implementation from the ParGeo library on a 20-core CPU, the full pipeline achieves speedups of up to 210× for uniform point distributions while degrading gracefully to approximately 1× in the worst case (all points on a sphere surface). The authors additionally report that GPU variants consume roughly 75× less energy than the CPU baseline for the favorable distribution (2601.19647).
The work extends a prior 2D GPU filtering study by Carrasco et al., which reported up to 160× end-to-end speedup over CGAL for uniform distributions. The 3D setting is harder: interior-point classification requires polyhedral containment tests rather than polygon tests, and candidate sets are typically larger relative to input size.
Prior acceleration strategies fall into three categories. Sequential filters include quickhull's initial quadrilateral filter, Skala et al.'s space-subdivision approach, Alshamrani et al.'s eight-vertex polygon with priority queues (up to 77× over Graham scan), and Ferrada et al.'s heaphull using Manhattan-distance metrics (1.7×–10× over CGAL). Parallel CPU approaches are represented by ParGeo's Pseudohull, which achieves up to 43.7× parallel speedup over Qhull and serves as this paper's primary baseline. GPU strategies include Stein et al.'s CudaHull (30× over Qhull) and Mei's rotational preprocessing (6×). Notably, several cited works lack publicly available implementations, which limits direct comparability; the authors benchmark instead against ParGeo as the strongest reproducible multicore baseline.
A key observation motivating the design is that RT cores provide hardware BVH traversal at 1×0 and ray/primitive intersection at 1×1, and prior work has successfully repurposed RT cores for non-graphical tasks such as neighbor search, range-minimum queries, and particle simulations.
Algorithm design
The pipeline consists of six phases, five executed on GPU and one (final hull construction) on CPU:
- Axis extreme points: six min/max reductions along 1×2, 1×3, 1×4 using grid-stride tree-based shared-memory reduction in 1×5.
- Corner points: for each of the 8 bounding-box corners, the nearest input point under Manhattan distance 1×6 is found via parallel reduction, yielding 8 additional vertices.
- Filtering polyhedron: the 14 points define a 24-triangle-face polyhedron inscribed within the true hull.
- Ray-based filtering: using OptiX, one ray is launched per input point from the point toward an interior reference point; any-hit shaders mark contained points as non-candidates, miss shaders mark candidates. The BVH build takes a constant ~1.9 ms regardless of input size since it contains only 24 triangles.
- Compaction: a three-level prefix scan computes scatter addresses — TC MMA operations handle 256-element segments per warp, followed by warp/block-level scans, a CUB-based block-total scan, and a downsweep kernel.
- Hull computation: the compacted set feeds Pseudohull.
Two implementation variants are evaluated: RTX (RT + TC + CUDA cores) and CUDA (TC + CUDA cores only, replacing BVH intersection with a direct loop over the 24 faces).
Experimental results
Experiments ran on an Intel Core Ultra 7 265K (8 performance + 12 efficiency cores) with an RTX 4090, using FP32 arithmetic, inputs from 1×7 to 1×8 points, multiple seeds, and 20–100 repetitions per measurement.
Filter phase: RTX reaches 1×9–75×0 speedup over 20-core Pseudohull on uniform distributions and 75×1–75×2 on sphere distributions. A notable crossover emerges: the CUDA variant outperforms RTX when the polyhedron is small, because it avoids BVH construction entirely, but its cost scales poorly with face count whereas hardware BVH traversal scales favorably.
End-to-end hull: combining the RTX filter with Pseudohull yields approximately 75×3 speedup over unfiltered 20-core Pseudohull for uniform distributions. For the sphere worst case (75×4), no points are filtered yet overall performance remains at ~75×5 — the filter adds negligible overhead. Sensitivity analysis over the displacement parameter 75×6 shows the filter beats unfiltered Quickhull at 75×7, beats the CPU filter at 75×8, and reaches optimal performance around 75×9. This is a strong practical claim: a mere 1% radial perturbation suffices for the GPU filter to be net-beneficial.
Energy: for uniform distributions, GPU implementations process ~62,000 points/Joule (~210 J total) versus ~830 points/Joule (~16,000 J) for the CPU. Average power draw is higher for GPUs (~280 W RTX, ~320 W CUDA vs ~140 W CPU), but execution duration is far shorter, so total energy is ~160×0 lower. In the sphere case all implementations converge to high consumption (~52,000 J), since 99.5% of time and energy is spent in the CPU-side hull computation itself — meaning the energy advantage is contingent on effective filtering.
Scalability: varying polyhedron complexity from 24 faces upward, the CUDA variant wins below roughly millisecond-scale differences while RTX exhibits exponentially better scaling on the logarithmic plot. Cross-architecture tests on Ampere (A100), Lovelace (RTX 4090), and Blackwell (RTX PRO 6000) show consistent or improving RTX-filter performance across generations, though Ampere lacks dedicated RT silicon and emulates traversal on CUDA cores.
An appendix profiling Pseudohull finds a counterintuitive result: 12 efficiency cores outperform 8 performance cores, indicating the algorithm benefits from parallelism width over clock speed, while efficiency-core configurations consume 25% less power and 50–75% less total energy than performance-core configurations.
Limitations and open questions
Several constraints qualify the reported results. The final hull computation still runs on the CPU, so in the worst-case sphere distribution the GPU contributes almost nothing and energy consumption matches the CPU baseline; the claimed energy efficiency is therefore distribution-dependent. The corner-selection heuristic uses Manhattan distance to bounding-box corners rather than maximum-volume tetrahedra, so the filtering polyhedron is not guaranteed to approximate the hull tightly. The scaling experiments vary face count but exclude the cost of constructing more robust polyhedra, which the authors state "considerably increases execution time." The evaluation covers only two synthetic distributions (uniform and sphere); behavior on real-world or adversarial data is not characterized. Finally, the paper leaves open whether a fully recursive GPU-resident version with dynamic BVH updates can push the filtering surface closer to the true hull without paying prohibitive construction costs.
Conclusion
The paper demonstrates that repurposing RT cores for geometric containment queries, combined with TC-accelerated stream compaction, yields substantial speedups (160×1 filter-phase, up to 160×2 end-to-end) and large energy savings over state-of-the-art multicore CPU implementations, with graceful degradation in adversarial cases. The main open question is whether dynamic, recursive BVH-based refinement on GPU can extend these gains to distributions where the initial 24-face polyhedron filters too few points.