Papers
Topics
Authors
Recent
Search
2000 character limit reached

CuACD: A Fully GPU-Resident Approximate Convex Decomposition

Published 23 Sep 2026 in cs.CG and cs.GR | (2609.28731v1)

Abstract: Approximate convex decomposition (ACD) converts triangle meshes into small sets of convex parts and is a standard preprocessing step for physics simulation, collision detection, and large-scale robot learning. The majority of modern ACD methods produce high-quality decompositions through an expensive search over candidate cutting planes, with per-mesh runtimes of tens of seconds that force game pipelines into overnight bakes and keep articulated-object datasets on CPU clusters for days. Prior work has accelerated isolated stages, most recently VisACD's GPU-based visibility metric, yet the dominant costs -- search, mesh cutting, and convex hull construction -- have remained on the CPU because their natural decomposition into many small homogeneous phases trails off in a fading last wave at every kernel boundary, and the variable-sized output of each phase forces a host round trip simply to allocate the next launch's input. We address these obstacles by adopting the warp, rather than the thread or thread block, as the unit of algorithm design, an idea introduced in the graph-processing community for a different pathology and which we adapt here to fuse the many heterogeneous phases of a computational-geometry pipeline into single warp-resident kernels, paired with a device-side heap allocator that lets the buffers between fused phases be sized and allocated on the device. Building on this template, we present CuACD (CUDA ACD), the first fully GPU-resident ACD system, together with a suite of reusable GPU components, released as open-source standalone CUDA modules that drop into any search-based ACD pipeline. On the V-HACD benchmark, PartNet-Mobility, and an Objaverse subset, CuACD achieves more than an order of magnitude of speedup over CoACD at matched or better quality.

Summary

  • The paper introduces CuACD, a fully GPU-resident approach to approximate convex decomposition (ACD), accelerateing the process by 78- to 104-fold compared to CoACD and 40- to 64-fold compared to VisACD on key benchmarks, with concavity measures remaining comparable or better and with no increase in part count.
  • CuACD enhances performance by using a warp-centric execution model, which efficiently handles irregular geometric operations and leverages co-operative reduction and symmetric paired work to optimize GPU utilization. At the same time, it improves memory management with custom GPU-side caching acceptable for most applications purposes.
  • Key contributions include accelerated mesh cutting, cheaper evaluations, and optimal memory management. It maintains concavity metrics of residual volume and bidirectional Hausdorff distance while using warf-based clamping of surface boundaries and significantly cuts time when compared to cpu implementations.

Approximate convex decomposition (ACD) is a preprocessing operation that replaces a triangle mesh with a collection of convex polyhedra suitable for collision detection. The practical importance of ACD follows from the efficiency of convex–convex proximity and intersection algorithms, whereas direct collision queries on arbitrary triangle meshes are substantially more expensive and less predictable. Exact convex decomposition is computationally intractable in general and commonly produces too many pieces for simulation, so modern systems optimize an approximate concavity objective under an implicit or explicit part-count constraint.

CuACD addresses the dominant cost of search-based ACD: the repeated evaluation of candidate cutting planes. Its central claim is that ACD is sufficiently coarse-grained across active parts, but sufficiently irregular within each candidate evaluation, to benefit from a warp-centric rather than thread-centric GPU design. The resulting system retains the search formulation and collision-aware concavity metric of CoACD while moving candidate generation, look-ahead search, mesh clipping, convex-hull construction, Hausdorff evaluation, connected-component analysis, and hull merging into GPU-resident execution (2609.28731).

Problem formulation and system objective

CuACD follows the recursive structure of CoACD. A worklist initially contains the input mesh. Parts whose concavity exceeds a threshold τ\tau are selected for further subdivision. For each selected part, the algorithm proposes cutting planes, evaluates them with a bounded look-ahead search, accepts the lowest-cost cut, and reinserts the resulting sub-parts. Once all parts satisfy the termination criterion, a greedy post-processing pass merges pairs whose joint hull remains below the threshold.

The concavity measure combines residual volume and bidirectional Hausdorff distance:

cost(P)=max⁡(kRv(P),dH(P,H(P))),\mathrm{cost}(P)=\max\left(kR_v(P),d_H(P,H(P))\right),

where H(P)H(P) is the convex hull of part PP, RvR_v is the radius of a sphere having the residual volume between PP and its hull, and kk calibrates the volume term. CuACD uses the cheaper residual-volume surrogate during look-ahead and evaluates the full metric for termination. This preserves the collision-aware motivation of CoACD: residual volume captures volumetric deviation, while Hausdorff distance penalizes large local surface errors.

The system proposes both axis-aligned planes and planes associated with concave edges. At the default configuration, a part receives 30 axis-aligned root candidates and up to 16 sampled concave edges, with four planes generated per edge. Thus, the nominal root pool contains up to 94 candidates. Deeper search uses a branching factor of five and depth two. This candidate enrichment is important when interpreting the quality results: part of CuACD's advantage over stock CoACD comes from a broader proposal distribution rather than GPU execution alone.

Warp-centric execution

The paper's main systems contribution is to make a warp—the 32 lanes executing in SIMT lockstep—the primary algorithmic unit. Conventional GPU implementations often assign individual threads or blocks to geometric operations and divide the pipeline into successive kernel launches. CuACD argues that this mapping is poorly matched to ACD for two reasons.

First, ACD exposes hundreds or thousands of independent parts and candidate evaluations, but individual tasks have highly variable workloads. A phase therefore ends with a long tail of underutilized threads and warps. Second, intermediate structures such as clipped vertices, boundary edges, hull faces, and BVH scratch buffers have data-dependent sizes. A phase-by-phase implementation must determine those sizes before launching the next phase, commonly forcing synchronization and host intervention.

CuACD instead uses two concurrency levels. At the coarse level, independent parts or candidate evaluations are assigned to resident warps. At the fine level, each warp cooperatively processes one irregular operation using warp votes, shuffles, reductions, filtering, and compact data structures. The paper's premise is that ACD's active worklist naturally supplies roughly the 10210^2 independent tasks needed to occupy a GPU at warp granularity, whereas a thread-granular design would require approximately 10410^4 independent units of work.

The warp-centric programming model is instantiated through four recurring patterns:

  • Cooperative traversal and reduction: lanes scan arrays at stride 32 and combine results using shuffle-based reductions.
  • Sorting and set construction: warp-level quick-sort supports dynamically sized arrays, deduplication, and membership operations.
  • Symmetric paired work: 16 two-lane groups process related entities such as edge endpoints or linked-list neighbors.
  • Divide and conquer: lanes or lane pairs independently process subproblems before warp-level merging.

This design does not eliminate irregular control flow. Rather, it confines irregular operations to a warp and exploits SIMD cooperation around them. The paper's approach is consequently most effective when there are many independent small or medium-sized geometry tasks, not when a single operation exposes little parallelism.

Device-resident memory management

The second enabling mechanism is a device-side heap allocator for variable-size intermediate buffers. CUDA's built-in device allocation primitives are unsuitable for CuACD because global allocation serialization becomes dominant at the allocation rate induced by repeated clipping and hull construction. Existing GPU allocators are also not sufficient for the system's variable-size mesh buffers because fragmentation and coalescing behavior are important.

CuACD divides the heap into 64 independent arenas. A warp is assigned to an arena according to its identifier, reducing contention among concurrent tasks. Each arena maintains power-of-two bins subdivided into 64 sub-bins, with a bitmap identifying nonempty lists. Boundary tags permit constant-time coalescing on free. The allocator therefore supports dynamic intermediate structures without requiring the host to inspect output sizes or relaunch kernels with newly computed capacities.

In the reported stress test, the allocator is approximately five times faster than CUDA allocation under contention-free conditions and 10–11 times faster when arenas are oversubscribed. These measurements support the claim that allocation is not merely an implementation detail: without a scalable device allocator, the fully resident pipeline would transfer its synchronization bottleneck into memory management.

The design also imposes a practical constraint. CuACD reserves 70% of free VRAM as a persistent heap pool at context creation. This simplifies allocation and stabilizes repeated workloads, but it reduces flexibility for applications sharing the GPU with other consumers. The heap does not compact memory across arenas or return freed memory to the system allocator, leaving fragmentation behavior under long-lived or heterogeneous workloads as an unresolved systems issue.

GPU mesh cutting and convex hull construction

Mesh-plane clipping is implemented as a sequence of warp-level filtering, sorting, deduplication, and traversal operations. The clipper classifies vertices by signed plane distance, identifies crossing edges, sorts edge keys, inserts unique intersection vertices, splits triangles into positive and negative submeshes, and reconstructs the cut boundary. The boundary is recovered from unpaired edge keys in the smaller output component. The cut is then capped using loop chaining, loop classification, hole bridging, and ear clipping.

The use of sorting to recover boundary edges is particularly appropriate for GPU execution. Interior edges appear twice in the edge multiset, while boundary edges appear once. Sorting transforms a topological boundary-identification problem into a local neighbor comparison. The subsequent cap construction remains more sequential: lane 0 maintains a cursor over a linked polygon representation, while the rest of the warp parallelizes point-in-triangle tests for candidate ears. This is a compromise between geometric correctness and SIMT efficiency rather than a fully parallel triangulation algorithm.

Convex hull construction is based on a warp-cooperative divide-and-conquer implementation of the Preparata–Hong algorithm using integer predicates. Input points are sorted along the longest bounding-box axis, partitioned into 16 chunks, and processed by two-lane hull groups. Partial hulls are merged through a gift-wrapping procedure in which the two lanes search from bridge endpoints and exchange state through shuffle operations.

For large point sets, CuACD adds a direction-extreme prefilter. Forty scans over directions sampled from a level-2 icosphere retain 80 extreme witnesses; a preliminary hull identifies points that cannot be extreme, after which the exact divide-and-conquer hull runs on the survivors. On V-HACD-style data with 65,536 points, this prefilter removes approximately 70% of points and produces more than a 12-fold speedup over the unfiltered GPU hull path. The prefilter is disabled below 1,024 points because its fixed scanning cost outweighs its savings.

The component benchmark reports substantial gains over Bullet's CPU hull implementation. For batches of 256 point sets with 65,536 points, the hybrid GPU method requires 20.53 ms for uniform-cube inputs and 17.61 ms for Gaussian inputs, compared with 6,839.78 ms and 5,561.01 ms for the CPU implementation. These results demonstrate that the relevant workload is not a single large hull but a large batch of independent hulls. That distinction separates CuACD from GPU hull algorithms optimized primarily for one large point cloud.

End-to-end performance

CuACD is evaluated on the 61-mesh V-HACD benchmark, 14,085 merged link meshes from PartNet-Mobility, and a 1,000-mesh Objaverse subset. Inputs are preprocessed using PaMO at a fixed configuration, and the preprocessing stage is excluded from timing. Baselines are threshold-swept to approximately match CoACD's default mean concavity of 0.05.

Benchmark Method Mean concavity Parts Time per mesh
V-HACD CoACD 0.0495 40.5 18.03 s
V-HACD VisACD 0.0604 37.5 9.60 s
V-HACD CuACD 0.0488 33.6 0.23 s
PartNet-Mobility CoACD 0.0465 21.1 12.82 s
PartNet-Mobility VisACD 0.0511 22.8 7.42 s
PartNet-Mobility CuACD 0.0458 21.0 0.16 s
Objaverse subset CoACD 0.0540 59.0 25.93 s
Objaverse subset VisACD 0.0702 59.9 15.92 s
Objaverse subset CuACD 0.0496 48.9 0.25 s

Relative to CoACD, the reported speedups are 78 times on V-HACD, 80 times on PartNet-Mobility, and 104 times on the Objaverse subset. CuACD is also reported to be 40–64 times faster than VisACD, the fastest prior GPU-assisted baseline in the comparison. On an RTX 3080 Mobile, the V-HACD mean is 0.64 seconds per mesh, still more than an order of magnitude faster than the CPU baselines.

The quality comparison is favorable but requires careful attribution. CuACD produces fewer parts and equal or lower mean concavity than the baselines in the matched operating points. However, its candidate pool is richer than CoACD's: CuACD uses 94 root candidates, including concave-edge planes, whereas the original CoACD configuration uses 60 axis-aligned candidates. A CPU implementation incorporating CuACD's concave-edge proposals and related hyperparameters produces 33.7 parts at mean maximum concavity 0.0481 and takes 20.85 seconds per mesh. This closely matches CuACD's 33.6 parts, 0.0488 concavity, and 0.23-second runtime.

That controlled comparison supports a narrower and stronger conclusion: the quality improvement is primarily attributable to candidate generation, while the large runtime improvement is primarily attributable to GPU-resident execution and warp-centric system design. The CPU port also shows that CuACD's search simplification does not inherently provide the speedup; the device implementation is the decisive factor.

Sensitivity, scaling, and application-level behavior

The search ablations indicate diminishing returns from deeper look-ahead. Increasing the depth from two to three lowers the mean part count from 33.57 to 32.89 but increases runtime from 0.23 to 0.41 seconds. Increasing the deeper branching factor from five to 15 raises runtime to 0.44 seconds while reducing the mean part count only to 33.07. Similarly, reducing the concave-edge iteration budget from 10 to three lowers runtime by approximately 17% but increases the mean part count from 33.57 to 35.20. These results justify the default operating point as a throughput-oriented compromise, although they do not establish that it is optimal for applications with a different cost function or part-count constraint.

Threshold sensitivity is comparatively favorable. Tightening τ\tau from 0.05 to 0.01 increases the mean part count from approximately 34 to 194 and raises runtime from approximately 0.23 to 0.5 seconds, while peak memory increases only from 1.4 to 1.6 GiB. Thus, within the tested range, a substantially finer decomposition costs roughly twice the runtime rather than an order-of-magnitude increase. The result is useful for pipelines that need to trade collision fidelity against collider complexity at runtime or preprocessing time.

The paper also reports scaling to inputs approaching one million triangles, with runtimes remaining in the multi-second range and memory within the 24 GiB capacity of the RTX 4090. A scene-scale experiment decomposes the 1,043,077-triangle Amazon Lumberyard Bistro interior at cost(P)=max⁡(kRv(P),dH(P,H(P))),\mathrm{cost}(P)=\max\left(kR_v(P),d_H(P,H(P))\right),0 in 18.73 seconds using 2,355.6 MiB of peak GPU memory. These measurements demonstrate that the implementation can process a complete large scene, but they should not be conflated with the per-mesh benchmark: the scene experiment uses a substantially tighter threshold and a different workload composition.

Limitations and open questions

CuACD retains hand-designed search heuristics inherited from the CoACD family. The paper explicitly concedes that some selected cuts may be suboptimal relative to learned proposal policies. The system therefore establishes a fast implementation of a classical search strategy, not a generally optimal decomposition procedure. Its comparison also excludes contemporary learning-based methods without released code, weights, or training data, so the reported quality and speed conclusions apply primarily to reproducible classical and GPU-assisted baselines.

The measured GPU occupancy is only 25–50%, limited by register pressure from intermediate geometric state. This does not prevent high throughput in the reported workloads, but it indicates that the implementation is not approaching hardware saturation in the usual occupancy sense. The paper leaves open whether lower register footprints, alternative state layouts, or more aggressive task scheduling would yield further gains without reducing geometric robustness.

The allocator's arena-local structure avoids global contention but does not compact memory across arenas. Long-running applications with changing mesh sizes could therefore exhibit fragmentation patterns not represented by the benchmark. The use of a persistent 70%-of-free-VRAM pool also presumes relatively exclusive GPU ownership. Finally, the experiments exclude PaMO preprocessing, despite its practical relevance in an end-to-end asset pipeline. Consequently, the reported speedups characterize decomposition after standardized preprocessing rather than total mesh-to-collider latency.

Conclusion

CuACD presents a fully GPU-resident implementation of search-based approximate convex decomposition. Its contribution is primarily architectural: warp-level task design handles irregular small-scale geometry operations, while a device-side coalescing allocator removes host synchronization caused by variable-size intermediate data. The resulting implementation reports 78–104-fold speedups over CoACD and 40–64-fold speedups over VisACD on three benchmarks, with comparable or better concavity and no increase in mean part count.

The controlled CPU comparison separates the system's contributions: enriched concave-edge proposals account for much of the quality improvement, whereas warp-centric fusion, GPU-resident allocation, and GPU implementations of clipping and hull construction account for the dominant performance gain. The remaining questions concern robustness under long-lived memory fragmentation, register-pressure reduction, end-to-end preprocessing cost, and the interaction between GPU-resident classical search and learned cut proposals.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Explain it Like I'm 14

1. What is this paper about?

This paper introduces CuACD, a new computer program for quickly breaking complicated 3D shapes into smaller, simpler shapes.

The method is called approximate convex decomposition (ACD). It takes a shape made from many triangles—such as a car, robot arm, chair, or game character—and divides it into several convex pieces.

A convex shape is like a ball, box, or solid lump with no deep dents. These shapes are useful because computers can quickly detect when two convex objects touch or collide.

The main idea of the paper is to move the entire decomposition process from the CPU to the GPU. GPUs are designed to perform many calculations at the same time, so this can make the process much faster.

2. What questions are the researchers asking?

The researchers are mainly trying to answer these questions:

  • Can the whole ACD process run on a GPU instead of repeatedly switching between the GPU and CPU?
  • Can a GPU handle difficult geometry tasks, such as cutting meshes and building convex hulls?
  • Can the new system produce results that are just as good as, or better than, existing methods?
  • How much faster can CuACD be when processing many 3D models?

The goal is not simply to make a rough answer quickly. The system should still create useful pieces that preserve important features of the original object, such as handles, wheels, corners, and joints.

3. How does the method work?

Breaking a shape into convex pieces

Imagine trying to turn a complicated clay sculpture into several simple blocks. CuACD repeatedly looks for a good flat cut that divides one complicated piece into two smaller pieces.

It continues cutting pieces that are too “dent-like” or non-convex. When all remaining pieces are simple enough, the process stops.

The system follows these general steps:

  1. Start with the complete 3D mesh.
  2. Find pieces that are too concave.
  3. Suggest many possible flat cuts.
  4. Test which cut produces the best result.
  5. Split the piece.
  6. Repeat until the pieces are simple enough.
  7. Try joining some pieces back together to avoid having too many pieces.

Measuring how complicated a piece is

The paper uses a measure called concavity. In everyday terms, this asks:

How different is this piece from the smallest convex shape that could cover it?

For example, if a shape has a deep hollow or large dent, its convex hull would fill in that dent. The amount of extra space and the distance between the original surface and the filled-in surface help measure its concavity.

The method uses two main measurements:

  • Volume difference: How much empty space is added when the dents are filled in.
  • Surface distance: How far the original surface is from the convex hull.

A piece with a smaller score is closer to being convex.

Why the GPU is difficult to use

GPUs work best when many tasks follow the same instructions. However, geometry processing often involves irregular decisions. For example, different parts may have different numbers of points, edges, or triangles.

The researchers solve this by using a GPU unit called a warp. A warp is a group of 32 GPU workers that perform instructions together.

Instead of assigning one tiny worker to each task, CuACD assigns one warp to each geometry task. The 32 workers in that warp cooperate, rather like a team of 32 people solving one small puzzle.

This approach is useful because CuACD often has hundreds or thousands of small geometry tasks happening at the same time.

Keeping everything on the GPU

Older approaches often had to send information back to the CPU between stages. This is similar to a factory that must stop production every time it needs to ask a manager how much storage space is available.

CuACD avoids this using a GPU memory allocator, also called a device-side heap allocator. It lets the GPU request and release memory by itself while the program is running.

As a result, the entire process—including searching for cuts, cutting meshes, building convex hulls, and measuring concavity—can stay on the GPU until the final answer is ready.

Important GPU components

The researchers create GPU versions of several geometry operations:

  • Cutting a mesh with a plane
  • Sorting and removing duplicate points
  • Building convex hulls
  • Finding connected parts of a mesh
  • Measuring distances between surfaces
  • Allocating temporary memory
  • Searching through possible cuts

These components are also released as separate CUDA modules, so other programmers may be able to reuse them.

4. What did the researchers find?

The experiments used three groups of 3D models:

  • The standard V-HACD benchmark
  • Robot and household-object models from PartNet-Mobility
  • A sample of models from Objaverse

The researchers compared CuACD with methods such as CoACD, NavACD, and VisACD.

Much faster processing

CuACD was reported to be dramatically faster than the comparison methods.

Dataset CuACD average time CoACD average time Approximate speedup
V-HACD 0.23 seconds 18.03 seconds 78×
PartNet-Mobility 0.16 seconds 12.82 seconds 80×
Objaverse subset 0.25 seconds 25.93 seconds 104×

In other words, a task that might take around 18 seconds with CoACD took about a quarter of a second with CuACD on the tested hardware.

The paper also reports that CuACD was about 40 to 64 times faster than VisACD, which was the fastest earlier GPU-assisted method included in the comparison.

Similar or better quality

Speed was not the only improvement. CuACD also produced:

  • Similar or lower concavity scores
  • Similar or fewer final pieces
  • Cuts that often followed meaningful features, such as handles, wheels, and joints

For example, on the V-HACD benchmark:

  • CoACD created about 40.5 pieces per model.
  • CuACD created about 33.6 pieces.
  • CuACD had a slightly lower average concavity score.

This suggests that the program was not simply making rougher results to become faster.

More detailed shapes are possible

The user can choose a stricter concavity limit. A stricter limit creates more pieces but preserves more detail.

According to the paper, lowering the threshold from 0.05 to 0.01 increased the average number of pieces from about 34 to about 194. Even so, the average processing time increased from around 0.23 seconds to about 0.5 seconds.

This is useful because users can choose between:

  • Fewer, simpler pieces for faster collision checking
  • More pieces for a closer representation of the original model

Strong performance on large inputs

The researchers also tested meshes with up to about one million triangles. CuACD continued to work within the memory limits of a single NVIDIA RTX 4090 GPU, although larger models naturally took longer.

5. Why are these results important?

Physics programs, video games, and robot simulators often need collision shapes before they can use 3D objects. Creating these shapes can take a long time, especially when thousands or millions of objects must be processed.

CuACD could make this preparation much faster. Possible benefits include:

  • Faster loading and preparation of game assets
  • Quicker robot simulation experiments
  • Faster creation of collision models for large 3D datasets
  • More interactive tools for artists and designers
  • Less need to process models overnight on large CPU computer clusters

For example, a robotics researcher working with thousands of objects could prepare a dataset in much less time. A game designer might also be able to generate collision shapes while working instead of waiting for an offline processing step.

Conclusion

This paper presents CuACD, a GPU-based system for dividing complicated 3D models into useful convex pieces. Its main innovation is that it keeps the entire process on the GPU and organizes the work around groups of 32 GPU workers called warps.

The experiments show very large speed improvements—often more than 70 times faster than earlier CPU-based methods—while maintaining or improving the quality of the decomposition.

The wider lesson is that even complicated, irregular geometry problems can sometimes be redesigned to run efficiently on GPUs. If the method becomes widely adopted, it could make physics simulation, robotics, video games, and large-scale 3D processing much faster and more convenient.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • Hardware generalizability is unresolved: the end-to-end evaluation is centered on an RTX 4090, with only a limited laptop RTX 3080 result; performance across different GPU architectures, memory capacities, CUDA versions, and non-NVIDIA accelerators is not established.
  • Multi-GPU and distributed scaling are unexplored: the paper does not determine whether independent meshes, large individual meshes, or decomposition worklists can be efficiently distributed across multiple GPUs.
  • CPU–GPU transfer costs are excluded from the main timing claims: the reported timings begin after preprocessing and mesh upload, so end-to-end latency from asset loading through collider export remains unknown.
  • The impact of PaMO preprocessing is not fully characterized: all principal comparisons use PaMO-processed meshes and exclude its runtime, but changes in topology, thin structures, curvature, and collision-relevant details caused by preprocessing are not evaluated.
  • Baseline comparisons are not fully controlled across algorithms: CuACD uses a different candidate-generation strategy and search procedure from CoACD, while the reported speedup combines GPU implementation, enriched candidate pools, and algorithmic changes; the isolated contribution of each change is not completely quantified.
  • Comparisons with learning-based ACD methods are absent: RL-ACD and other learned approaches are excluded because their artifacts are unavailable, leaving the relative quality, speed, and generalization of CuACD versus current learned methods unresolved.
  • Downstream physics performance is not measured: the paper reports concavity and number of parts, but does not evaluate collision-detection throughput, contact stability, penetration errors, simulation accuracy, or training performance in Bullet, PhysX, MuJoCo, or robotics simulators.
  • The relationship between the selected concavity metric and practical collision quality remains uncertain: it is not shown whether lower bidirectional Hausdorff/volume-based cost consistently produces better contact behavior for diverse physical tasks.
  • Hausdorff evaluation is itself approximate: distances are estimated from uniformly sampled points, but the paper does not quantify sampling error, missed worst-case deviations, or sensitivity to mesh scale and triangle density.
  • The use of the cheap kRvkR_v surrogate during tree search may alter search quality: the frequency with which surrogate-optimal cuts differ from cuts selected using the full Hausdorff-based objective is not reported.
  • Quality evaluation relies primarily on aggregate means: mean concavity and mean part count can hide severe outliers; per-mesh distributions, worst-case errors, percentile runtimes, and failure-prone geometry categories are not analyzed in detail.
  • The paper does not report systematic failure cases: robustness on non-manifold meshes, self-intersections, open boundaries, inverted normals, duplicate vertices, extreme coordinate ranges, very thin shells, and highly disconnected or nested components remains unclear.
  • The stated watertight clipping assumptions limit applicability: boundary recovery depends on manifold edge multiplicities, but the behavior when those assumptions fail is not specified or experimentally evaluated.
  • Numerical robustness of the clipping and cap triangulation stages is insufficiently characterized: nearly coplanar vertices, repeated intersections, degenerate cap polygons, narrow sliver triangles, and precision issues in plane-side tests may cause incorrect topology.
  • The integer-coordinate convex-hull formulation is not fully specified: the conversion from input floating-point coordinates to integer predicates, including scaling, overflow limits, quantization error, and behavior over large coordinate ranges, is not evaluated.
  • The exactness claim for convex hull construction is incomplete: although integer determinants avoid floating-point tolerance decisions, the paper does not provide formal correctness guarantees or adversarial tests for degeneracies such as coplanar, collinear, duplicated, or cospherical points.
  • Allocator behavior under long and diverse workloads is unresolved: experiments use a particular stress-test distribution and a fixed 64-arena configuration, but sustained fragmentation, out-of-memory behavior, arena imbalance, allocation failures, and performance across larger candidate trees are not reported.
  • The fixed reservation of 70% of free VRAM may be impractical in shared or memory-constrained environments: its effect on other GPU workloads, CUDA context coexistence, and peak-memory variability is not studied.
  • Memory scaling for simultaneous large meshes is unclear: the million-triangle experiment uses a single GPU and reports aggregate behavior, but the maximum number of concurrently decomposed assets and the effect of batching heterogeneous meshes are not established.
  • Warp-level load imbalance is not comprehensively quantified: occupancy is reported as 25–50%, but the paper does not separate losses from branch divergence, unequal mesh sizes, candidate-size variance, register pressure, memory latency, and allocator contention.
  • The choice of warp as the universal task unit may be suboptimal for very small or very large parts: no comparison is provided against adaptive thread-, warp-, or block-level scheduling.
  • The quick-sort implementation lacks broad adversarial evaluation: pivot robustness, worst-case behavior, duplicate-heavy inputs, nearly sorted inputs, and comparisons with appropriately configured radix or segmented GPU sorting are not reported.
  • The direction-extreme hull prefilter may have distribution-dependent benefits: its fixed icosphere directions and threshold of 1,024 points are not justified across anisotropic, highly concave, adversarial, or nonuniform point distributions.
  • Search hyperparameters may not transfer across datasets: the selected values for candidate counts, look-ahead depth, branching factor, and concave-edge sampling are evaluated on the reported benchmarks, but no principled tuning method or cross-domain robustness analysis is given.
  • The threshold sweep does not establish monotonic quality guarantees: although tighter τ\tau generally produces more parts, the paper does not determine whether per-part concavity, collision quality, or part count is monotonic for every input.
  • The greedy hull-merge pass may be order-dependent: sensitivity to candidate ordering, merge ordering, and local minima is not analyzed, and no comparison with globally or jointly optimized merging is provided.
  • The effect of connected-component splitting on semantic or physical decomposition is unknown: splitting disconnected components may be beneficial computationally, but its consequences for articulated objects, intentionally connected collision structures, and downstream asset semantics are not evaluated.
  • Temporal and dynamic use cases remain unexplored: the method is evaluated for static preprocessing, while animated meshes, deforming geometry, topology changes, and incremental updates are not addressed.
  • Determinism and reproducibility are not established: atomic operations, best-effort union-find updates, parallel allocation, and scheduling may produce nondeterministic outputs or runtimes, but repeatability across runs is not reported.
  • Energy efficiency and total resource cost are omitted: the paper reports wall-clock speedups but does not compare GPU energy consumption, carbon cost, or total hardware cost against CPU clusters and hybrid GPU pipelines.
  • Interactive authoring benefits are asserted but not measured: whether sub-second decomposition improves artist iteration, reduces manual proxy creation, or meets interactive application latency requirements remains an open empirical question.
  • The reusable CUDA components are not evaluated outside CuACD: portability to other search-based ACD systems, mesh-processing pipelines, compiler configurations, and different data layouts is claimed but not demonstrated quantitatively.

Practical Applications

Immediate Applications

  • Accelerated collision-geometry preprocessing for game engines (Industry—gaming and interactive 3D)
    • Potential tools: a CUDA-based asset-import plugin, automatic collider generation in physics engines, or a command-line batch converter for engines such as Bullet, PhysX, MuJoCo, and game-editor toolchains.
    • Workflow: upload a mesh to the GPU → select a concavity threshold and part-count target → generate convex hulls → export collider geometry and metadata.
    • Dependencies: NVIDIA CUDA-compatible hardware, sufficient VRAM, and integration with the target engine’s convex-shape format. The paper’s timings exclude PaMO mesh preprocessing, so production pipelines may still need a separate simplification or remeshing step.
  • High-throughput collider generation for robotic simulation (Industry/Academia—robotics and embodied AI)
    • Potential products and workflows: automatic collider generation for Isaac Sim, MuJoCo, Bullet, ManiSkill, and differentiable robotics environments; dataset-preparation services that preprocess thousands of robot or household-object models.
    • Practical benefit: faster construction of simulation-ready assets and more frequent regeneration of colliders when meshes or decomposition thresholds change.
    • Dependencies: the generated convex parts must be validated for watertightness, contact stability, and simulator-specific limits on the number of convex bodies. GPU memory requirements may become significant for large articulated scenes or simultaneous batch processing.
  • Interactive 3D content-authoring assistance (Industry—CAD, digital content creation, and game development)
    • Potential tools: viewport overlays showing convex parts, automatic “generate collider” buttons, and sliders for coarse-to-fine decomposition.
    • Dependencies: latency will vary with mesh complexity, GPU model, asset batching, and required output quality. The paper reports averages; worst-case latency and user-interface integration are not evaluated.
  • Batch preprocessing for large 3D and robotics datasets (Academia/Industry—3D AI and simulation)
    • Potential workflow: distributed GPU workers read meshes, run CuACD, export convex decompositions, and cache them alongside dataset records.
    • Practical benefit: reduced CPU-cluster occupancy and faster preparation of training environments.
    • Dependencies: throughput at cluster scale, GPU scheduling, data-transfer overhead, and licensing or format compatibility of source datasets. The study evaluates a single GPU extensively but does not report multi-GPU orchestration.
  • Drop-in GPU acceleration for existing computational-geometry pipelines (Industry/Academia—software infrastructure)
    • Potential applications: GPU mesh slicing, collision preprocessing, geometric validation, remeshing support, proximity analysis, and batched convex-hull computation.
    • Potential software artifact: a CUDA library exposing kernels over structure-of-arrays mesh buffers, allowing existing GPU geometry systems to reuse the components without adopting the full CuACD search pipeline.
    • Dependencies: CUDA programming expertise, careful memory-management integration, and validation against numerical and topological edge cases. Performance benefits may be smaller for workloads lacking many independent small geometry tasks.
  • Faster collision-query preparation for physics-based animation (Industry—animation, virtual prototyping, and simulation)
    • Potential workflow: generate multiple collider levels of detail using different τ\tau values and select among them according to frame rate or simulation accuracy requirements.
    • Dependencies: the relationship between the paper’s concavity metric and application-specific contact quality must be validated for each physics engine. A lower concavity does not automatically guarantee better stability for friction, joints, stacking, or thin-shell contacts.
  • Teaching and research infrastructure for GPU computational geometry (Academia)
    • Potential uses: laboratory exercises implementing warp-level sorting, GPU clipping, convex hulls, union-find, and device-side allocation; comparative studies of block-centric versus warp-centric designs.
    • Dependencies: the code’s CUDA and GPU-generation assumptions, documentation quality, and the need to distinguish reproducible benchmark results from hardware-specific optimizations.
  • Automated collision-mesh quality control (Industry—asset validation and QA)
    • Potential workflow: every asset commit triggers GPU decomposition; the system reports part count, maximum concavity, memory use, and failed topology checks.
    • Dependencies: project-specific acceptance thresholds and robust handling of non-manifold, self-intersecting, or incorrectly scaled meshes. The paper reports zero failures on its evaluated inputs but does not establish universal robustness.

Long-Term Applications

  • Real-time or near-real-time dynamic collider generation (Long-term—games, robotics, and interactive simulation)
    • Potential products: runtime destruction systems, dynamic object spawning with automatic colliders, and adaptive collision geometry for deformable robots.
    • Dependencies: temporal coherence, incremental updates, bounded worst-case latency, and memory reclamation for frequently changing meshes. Recomputing a full decomposition for every frame is unlikely to be practical.
  • GPU-accelerated motion planning and navigation geometry (Long-term—robotics and game AI)
    • Potential workflow: generate convex parts for obstacles, combine them with navigation meshes or free-space constraints, and update the representation as environments change.
    • Dependencies: the current system optimizes collision-oriented concavity rather than guaranteeing navigational correctness. Additional methods are needed to preserve clearance, avoid blocking valid paths, and support dynamic scenes.
  • Physics-aware synthetic-data generation at scale (Long-term—robotics and machine learning)
    • Potential outputs: training corpora containing meshes, convex decompositions, contact labels, trajectories, and rendered observations.
    • Dependencies: decomposition quality must be correlated with downstream learning performance. Dataset bias, simulator fidelity, GPU availability, and robust handling of real-world scans remain open concerns.
  • Learned or adaptive decomposition systems (Long-term—AI and computational geometry)
    • Potential tools: learned cut proposal networks, per-object threshold predictors, or hybrid systems that use neural inference for proposal generation and exact GPU geometry for validation.
    • Dependencies: training data, generalization across object categories and mesh quality, reproducibility, and maintaining geometric guarantees. The paper does not test learned proposal mechanisms.
  • Multi-GPU and cloud-scale geometry services (Long-term—cloud computing and enterprise software)
    • Potential architecture: a queue of mesh-decomposition jobs, GPU workers, cached multilevel collider outputs, and automatic selection of τ\tau based on application requirements.
    • Dependencies: batching efficiency, GPU isolation, VRAM fragmentation, fault tolerance, data-transfer costs, and support for non-NVIDIA accelerators. The current implementation is CUDA-specific and does not demonstrate distributed execution.
  • Porting the warp-centric design to other irregular geometry problems (Long-term—GPU software and scientific computing)
    • Potential tools: a general-purpose warp-centric geometry framework with reusable allocators, sorting, clipping, hull, and traversal abstractions.
    • Dependencies: each target algorithm must expose enough independent work and tolerate SIMT control-flow divergence. Workloads with few large tasks, strong global dependencies, or highly skewed sizes may not achieve the reported gains.
  • Hardware-portable GPU geometry libraries (Long-term—systems research)
    • Potential outcome: portable libraries for convex decomposition and batched computational geometry in robotics, CAD, and simulation stacks.
    • Dependencies: warp-level intrinsics, memory-allocation primitives, integer-predicate behavior, and performance portability differ across platforms. Porting may require redesign rather than straightforward translation.
  • Formal quality and safety guarantees for physical simulation (Long-term—physics, robotics, and safety-critical engineering)
    • Potential applications: certified collider generation for autonomous robots, virtual prototyping, and safety-critical digital twins.
    • Dependencies: the paper demonstrates matched or improved geometric metrics and part counts, but it does not establish formal bounds on simulated dynamics. Such guarantees would require experiments across physics engines, contact models, scales, and adversarial geometries.

Glossary

  • Approximate convex decomposition (ACD): The process of partitioning a nonconvex mesh into a small number of approximately convex components. “Approximate convex decomposition (ACD) converts triangle meshes into small sets of convex parts”
  • AtomicCAS: An atomic compare-and-swap operation used to update shared memory safely in parallel. “we attempt the rank update with a single atomicCAS”
  • Bidirectional Hausdorff distance: A geometric distance measuring the greatest deviation between two shapes in both directions. “CoACD's concavity metric unchanged. For a part PP with convex hull H(P)H(P), the concavity is”
  • Bitonic network: A fixed comparison network used for parallel sorting. “we apply a bitonic network to sort them directly”
  • Bounding-volume hierarchy (BVH): A tree structure that groups geometric objects within nested bounding volumes to accelerate spatial queries. “the BVH scratch for Hausdorff queries”
  • Concavity objective: A numerical function measuring how much a shape differs from its convex hull. “an equally expensive concavity objective---a scalar measuring how far a piece deviates from convex”
  • Concave-edge plane: A cutting plane positioned using a mesh edge whose incident faces form a strongly concave angle. “Concave-edge planes target contact-relevant features”
  • Connected-component analysis: The identification of separate, mutually disconnected regions in a mesh or graph. “connected-component analysis, and Hausdorff distance”
  • Coalescing-binned allocator: A memory allocator that groups free blocks into size classes and merges adjacent free blocks. “The remedy is the coalescing-binned machinery”
  • Computational geometry: The study and development of algorithms for geometric objects and spatial relationships. “the irregular kernels of computational geometry”
  • Convex hull: The smallest convex shape containing all points of a given shape or point set. “building the convex hulls of the two resulting pieces”
  • Convex collider: A convex geometric representation used for efficient collision detection in simulation. “Modern physics engines rely on convex colliders”
  • CUDA: NVIDIA’s platform and programming model for general-purpose GPU computing. “We present CuACD (CUDA ACD)”
  • Device-resident: Stored and processed in GPU memory without transferring intermediate data to the host CPU. “The first fully GPU-resident ACD algorithm and system”
  • Dihedral angle: The angle between two planes, such as the planes of adjacent polygonal faces. “we sample up to $N_{\mathrm{ce}$ concave edges with dihedral angle ≥200°\ge 200\degree”
  • Ear-clipping: A polygon-triangulation method that repeatedly removes suitable triangular “ears” from a polygon. “Bridge holes by ear-clipping into cap CC”
  • Floating-point tolerance: A numerical threshold used to account for rounding errors in floating-point computations. “the kernel is free of floating-point tolerance decisions”
  • GPU-resident: Existing entirely in GPU memory and executing without intermediate host-CPU coordination. “a fully GPU-resident ACD algorithm”
  • Heap allocator: A memory-management component that dynamically allocates and frees storage during execution. “a device-side heap allocator”
  • Hybrid prefilter: A preliminary filtering step that removes points unlikely to contribute to the final convex hull before exact construction. “the hybrid variant discards them first”
  • Integer determinant: A determinant computed over integer coordinates, commonly used for exact geometric orientation tests. “All predicates are integer determinants”
  • Intersection-free remeshing: Mesh reconstruction that avoids introducing intersecting geometric elements. “parallel intersection-free remeshing and mesh simplification”
  • Kernel: A GPU function executed concurrently by many GPU threads or warps. “the dominant costs---search, mesh cutting, and convex hull construction---have remained on the CPU”
  • Lock-free: A concurrency property in which system-wide progress occurs without mutual-exclusion locks. “We implement this as a warp-parallel union-find over triangles”
  • Look-ahead tree search: A bounded recursive search that evaluates possible future actions before selecting the current action. “a parallel look-ahead tree search”
  • Manifold property: A geometric condition in which each local region has a topology resembling Euclidean space; for a closed triangle mesh, an interior edge is typically incident to two faces. “We use the manifold property---an interior edge is shared by exactly two triangles”
  • Monte Carlo tree search (MCTS): A randomized tree-search algorithm that estimates the quality of actions through repeated simulations or sampling. “drives the search with Monte Carlo tree search”
  • OptiX: NVIDIA’s GPU ray-tracing framework for accelerating ray-based geometric queries. “evaluated with OptiX ray queries on the GPU”
  • Plane clipping: The operation of splitting geometry against a plane and retaining the resulting portions. “Mesh--plane clipping”
  • Ray-casting: A geometric technique that shoots rays through a shape to determine containment or intersections. “Classify each LiL_i into outer/inner loop by 2D ray-casting”
  • SIMD: A parallel execution model in which one instruction operates on multiple data elements simultaneously. “each warp is a $32$-lane SIMD processor”
  • SIMT: A GPU execution model in which groups of threads execute a common instruction stream while operating on separate data. “The irregular, data-dependent control flow of convex hull construction”
  • Spatial indexing: Organizing geometric data to accelerate spatial searches and proximity queries. “rasterization, spatial indexing, hierarchy building, and traversal”
  • Struct-of-arrays: A data layout that stores each field of a collection in a separate contiguous array. “CuACD's kernels instead work on plain struct-of-arrays buffers”
  • Streaming multiprocessor (SM): A GPU hardware unit that schedules and executes groups of parallel threads. “a few hundred independent tasks across the GPU's streaming multiprocessors”
  • Tail of stragglers: The small remainder of active work that leaves many parallel processors idle near the end of a computation. “each kernel boundary ends in a tail of stragglers”
  • Temporal coherence: The exploitation of similarity between successive time steps in an animated or changing scene. “extend ACD to animated meshes via temporal coherence”
  • Union-find: A data structure supporting efficient maintenance of disjoint sets and connectivity queries. “a warp-parallel union-find over triangles”
  • Variable-sized buffer: A memory region whose required size is determined dynamically during execution. “the variable-size vertex arrays a plane cut emits”
  • Warp: A fixed group of GPU lanes that execute instructions together, typically as a single scheduling unit. “only the warp---a fixed group of $32$ lanes sharing one instruction pointer”
  • Warp-level intrinsic: A GPU instruction or primitive enabling communication or voting among lanes within a warp. “using warp-level vote and shuffle intrinsics”
  • Watertight mesh: A mesh with no gaps or cracks, so its surface forms a closed boundary. “This produces two watertight sub-meshes”
  • Voxelized recursive bisection: Repeatedly splitting a voxel representation of an object into smaller regions. “V-HACD replaces triangle-level clustering with a voxelized recursive bisection”
  • Worklist: A dynamically maintained collection of tasks or geometric parts awaiting processing. “The algorithm maintains a worklist of parts that still require further decomposition”

Tweets

Sign up for free to view the 2 tweets with 82 likes about this paper.