Papers
Topics
Authors
Recent
Search
2000 character limit reached

Convex Hull 3D Filtering with GPU Ray Tracing and Tensor Cores

Published 27 Jan 2026 in cs.CG and cs.DC | (2601.19647v1)

Abstract: In recent years, applications such as real-time simulations, autonomous systems, and video games increasingly demand the processing of complex geometric models under stringent time constraints. Traditional geometric algorithms, including the convex hull, are subject to these challenges. A common approach to improve performance is scaling computational resources, which often results in higher energy consumption. Given the growing global concern regarding sustainable use of energy, this becomes a critical limitation. This work presents a 3D preprocessing filter for the convex hull algorithm using ray tracing and tensor core technologies. The filter builds a delimiter polyhedron based on Manhattan distances that discards points from the original set. The filter is evaluated on two point distributions: uniform and sphere. Experimental results show that the proposed filter, combined with convex hull construction, accelerates the computation of the 3D convex hull by up to 200×200 \times with respect to a CPU parallel implementation. This research demonstrates that geometric algorithms can be accelerated through massive parallelism while maintaining efficient energy utilization. Beyond execution time and speedup evaluation, we also analyze GPU energy consumption, showing that the proposed preprocessing filter not only reduces the computational workload but also achieves performance gains with controlled energy usage. These results highlight the dual benefit of the method in terms of both speed and energy efficiency, reinforcing its applicability in modern high-performance scenarios.

Summary

  • The paper introduces a GPU-based processing pipeline for 3D convex hull filtering, leveraging CUDA, Tensor, and Ray Tracing cores. The method achieves up to $210 imes$ speed and $75 imes$ energy efficiency improvements over CPU for uniform point distributions; performance graphifies to floor for the worst case scenario (sphere).
  • The new approach relies on ray tracing for efficient containment checks and TC operations on prefix scans. The process involves several computational phases executed primarily on the GPU with exception for the final hull computation that remains on the CPU, though this is highlighted as a potential area for future optimization.
  • The workload of the RT cores greatly varies by the distribution shape of the input points. Uniform distributions are favorable for GPU-based methods. Real-world data potentially provides a challenge because this method is not well tested for this. The paper suggests that further development may push the cut-off curve closer to ideal with optimized GPU dynamic BHI budgeting strategies.

Overview

This paper presents a GPU-based preprocessing filter for the 3D convex hull problem that combines three heterogeneous compute resources available on modern NVIDIA GPUs: CUDA cores, Tensor Cores (TC), and Ray Tracing (RT) cores. The filter constructs a 24-face filtering polyhedron from 14 extreme points of the input set, discards interior points via hardware-accelerated ray/polygon intersection, and compacts the surviving candidate set using a TC-accelerated prefix scan. Evaluated against the Pseudohull implementation from the ParGeo library on a 20-core CPU, the full pipeline achieves speedups of up to 210×210\times for uniform point distributions while degrading gracefully to approximately 1×1\times in the worst case (all points on a sphere surface). The authors additionally report that GPU variants consume roughly 75×75\times less energy than the CPU baseline for the favorable distribution (2601.19647).

The work extends a prior 2D GPU filtering study by Carrasco et al., which reported up to 160×160\times end-to-end speedup over CGAL for uniform distributions. The 3D setting is harder: interior-point classification requires polyhedral containment tests rather than polygon tests, and candidate sets are typically larger relative to input size.

Prior acceleration strategies fall into three categories. Sequential filters include quickhull's initial quadrilateral filter, Skala et al.'s space-subdivision approach, Alshamrani et al.'s eight-vertex polygon with priority queues (up to 77×77\times over Graham scan), and Ferrada et al.'s heaphull using Manhattan-distance metrics (1.7×1.7\times–10×10\times over CGAL). Parallel CPU approaches are represented by ParGeo's Pseudohull, which achieves up to 43.7×43.7\times parallel speedup over Qhull and serves as this paper's primary baseline. GPU strategies include Stein et al.'s CudaHull (30×30\times over Qhull) and Mei's rotational preprocessing (6×6\times). Notably, several cited works lack publicly available implementations, which limits direct comparability; the authors benchmark instead against ParGeo as the strongest reproducible multicore baseline.

A key observation motivating the design is that RT cores provide hardware BVH traversal at 1×1\times0 and ray/primitive intersection at 1×1\times1, and prior work has successfully repurposed RT cores for non-graphical tasks such as neighbor search, range-minimum queries, and particle simulations.

Algorithm design

The pipeline consists of six phases, five executed on GPU and one (final hull construction) on CPU:

  1. Axis extreme points: six min/max reductions along 1×1\times2, 1×1\times3, 1×1\times4 using grid-stride tree-based shared-memory reduction in 1×1\times5.
  2. Corner points: for each of the 8 bounding-box corners, the nearest input point under Manhattan distance 1×1\times6 is found via parallel reduction, yielding 8 additional vertices.
  3. Filtering polyhedron: the 14 points define a 24-triangle-face polyhedron inscribed within the true hull.
  4. Ray-based filtering: using OptiX, one ray is launched per input point from the point toward an interior reference point; any-hit shaders mark contained points as non-candidates, miss shaders mark candidates. The BVH build takes a constant ~1.9 ms regardless of input size since it contains only 24 triangles.
  5. Compaction: a three-level prefix scan computes scatter addresses — TC MMA operations handle 256-element segments per warp, followed by warp/block-level scans, a CUB-based block-total scan, and a downsweep kernel.
  6. Hull computation: the compacted set feeds Pseudohull.

Two implementation variants are evaluated: RTX (RT + TC + CUDA cores) and CUDA (TC + CUDA cores only, replacing BVH intersection with a direct loop over the 24 faces).

Experimental results

Experiments ran on an Intel Core Ultra 7 265K (8 performance + 12 efficiency cores) with an RTX 4090, using FP32 arithmetic, inputs from 1×1\times7 to 1×1\times8 points, multiple seeds, and 20–100 repetitions per measurement.

Filter phase: RTX reaches 1×1\times9–75×75\times0 speedup over 20-core Pseudohull on uniform distributions and 75×75\times1–75×75\times2 on sphere distributions. A notable crossover emerges: the CUDA variant outperforms RTX when the polyhedron is small, because it avoids BVH construction entirely, but its cost scales poorly with face count whereas hardware BVH traversal scales favorably.

End-to-end hull: combining the RTX filter with Pseudohull yields approximately 75×75\times3 speedup over unfiltered 20-core Pseudohull for uniform distributions. For the sphere worst case (75×75\times4), no points are filtered yet overall performance remains at ~75×75\times5 — the filter adds negligible overhead. Sensitivity analysis over the displacement parameter 75×75\times6 shows the filter beats unfiltered Quickhull at 75×75\times7, beats the CPU filter at 75×75\times8, and reaches optimal performance around 75×75\times9. This is a strong practical claim: a mere 1% radial perturbation suffices for the GPU filter to be net-beneficial.

Energy: for uniform distributions, GPU implementations process ~62,000 points/Joule (~210 J total) versus ~830 points/Joule (~16,000 J) for the CPU. Average power draw is higher for GPUs (~280 W RTX, ~320 W CUDA vs ~140 W CPU), but execution duration is far shorter, so total energy is ~160×160\times0 lower. In the sphere case all implementations converge to high consumption (~52,000 J), since 99.5% of time and energy is spent in the CPU-side hull computation itself — meaning the energy advantage is contingent on effective filtering.

Scalability: varying polyhedron complexity from 24 faces upward, the CUDA variant wins below roughly millisecond-scale differences while RTX exhibits exponentially better scaling on the logarithmic plot. Cross-architecture tests on Ampere (A100), Lovelace (RTX 4090), and Blackwell (RTX PRO 6000) show consistent or improving RTX-filter performance across generations, though Ampere lacks dedicated RT silicon and emulates traversal on CUDA cores.

An appendix profiling Pseudohull finds a counterintuitive result: 12 efficiency cores outperform 8 performance cores, indicating the algorithm benefits from parallelism width over clock speed, while efficiency-core configurations consume 25% less power and 50–75% less total energy than performance-core configurations.

Limitations and open questions

Several constraints qualify the reported results. The final hull computation still runs on the CPU, so in the worst-case sphere distribution the GPU contributes almost nothing and energy consumption matches the CPU baseline; the claimed energy efficiency is therefore distribution-dependent. The corner-selection heuristic uses Manhattan distance to bounding-box corners rather than maximum-volume tetrahedra, so the filtering polyhedron is not guaranteed to approximate the hull tightly. The scaling experiments vary face count but exclude the cost of constructing more robust polyhedra, which the authors state "considerably increases execution time." The evaluation covers only two synthetic distributions (uniform and sphere); behavior on real-world or adversarial data is not characterized. Finally, the paper leaves open whether a fully recursive GPU-resident version with dynamic BVH updates can push the filtering surface closer to the true hull without paying prohibitive construction costs.

Conclusion

The paper demonstrates that repurposing RT cores for geometric containment queries, combined with TC-accelerated stream compaction, yields substantial speedups (160×160\times1 filter-phase, up to 160×160\times2 end-to-end) and large energy savings over state-of-the-art multicore CPU implementations, with graceful degradation in adversarial cases. The main open question is whether dynamic, recursive BVH-based refinement on GPU can extend these gains to distributions where the initial 24-face polyhedron filters too few points.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.