CuACD: A Fully GPU-Resident Approximate Convex Decomposition
Abstract: Approximate convex decomposition (ACD) converts triangle meshes into small sets of convex parts and is a standard preprocessing step for physics simulation, collision detection, and large-scale robot learning. The majority of modern ACD methods produce high-quality decompositions through an expensive search over candidate cutting planes, with per-mesh runtimes of tens of seconds that force game pipelines into overnight bakes and keep articulated-object datasets on CPU clusters for days. Prior work has accelerated isolated stages, most recently VisACD's GPU-based visibility metric, yet the dominant costs -- search, mesh cutting, and convex hull construction -- have remained on the CPU because their natural decomposition into many small homogeneous phases trails off in a fading last wave at every kernel boundary, and the variable-sized output of each phase forces a host round trip simply to allocate the next launch's input. We address these obstacles by adopting the warp, rather than the thread or thread block, as the unit of algorithm design, an idea introduced in the graph-processing community for a different pathology and which we adapt here to fuse the many heterogeneous phases of a computational-geometry pipeline into single warp-resident kernels, paired with a device-side heap allocator that lets the buffers between fused phases be sized and allocated on the device. Building on this template, we present CuACD (CUDA ACD), the first fully GPU-resident ACD system, together with a suite of reusable GPU components, released as open-source standalone CUDA modules that drop into any search-based ACD pipeline. On the V-HACD benchmark, PartNet-Mobility, and an Objaverse subset, CuACD achieves more than an order of magnitude of speedup over CoACD at matched or better quality.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper introduces CuACD, a new computer program for quickly breaking complicated 3D shapes into smaller, simpler shapes.
The method is called approximate convex decomposition (ACD). It takes a shape made from many triangles—such as a car, robot arm, chair, or game character—and divides it into several convex pieces.
A convex shape is like a ball, box, or solid lump with no deep dents. These shapes are useful because computers can quickly detect when two convex objects touch or collide.
The main idea of the paper is to move the entire decomposition process from the CPU to the GPU. GPUs are designed to perform many calculations at the same time, so this can make the process much faster.
2. What questions are the researchers asking?
The researchers are mainly trying to answer these questions:
- Can the whole ACD process run on a GPU instead of repeatedly switching between the GPU and CPU?
- Can a GPU handle difficult geometry tasks, such as cutting meshes and building convex hulls?
- Can the new system produce results that are just as good as, or better than, existing methods?
- How much faster can CuACD be when processing many 3D models?
The goal is not simply to make a rough answer quickly. The system should still create useful pieces that preserve important features of the original object, such as handles, wheels, corners, and joints.
3. How does the method work?
Breaking a shape into convex pieces
Imagine trying to turn a complicated clay sculpture into several simple blocks. CuACD repeatedly looks for a good flat cut that divides one complicated piece into two smaller pieces.
It continues cutting pieces that are too “dent-like” or non-convex. When all remaining pieces are simple enough, the process stops.
The system follows these general steps:
- Start with the complete 3D mesh.
- Find pieces that are too concave.
- Suggest many possible flat cuts.
- Test which cut produces the best result.
- Split the piece.
- Repeat until the pieces are simple enough.
- Try joining some pieces back together to avoid having too many pieces.
Measuring how complicated a piece is
The paper uses a measure called concavity. In everyday terms, this asks:
How different is this piece from the smallest convex shape that could cover it?
For example, if a shape has a deep hollow or large dent, its convex hull would fill in that dent. The amount of extra space and the distance between the original surface and the filled-in surface help measure its concavity.
The method uses two main measurements:
- Volume difference: How much empty space is added when the dents are filled in.
- Surface distance: How far the original surface is from the convex hull.
A piece with a smaller score is closer to being convex.
Why the GPU is difficult to use
GPUs work best when many tasks follow the same instructions. However, geometry processing often involves irregular decisions. For example, different parts may have different numbers of points, edges, or triangles.
The researchers solve this by using a GPU unit called a warp. A warp is a group of 32 GPU workers that perform instructions together.
Instead of assigning one tiny worker to each task, CuACD assigns one warp to each geometry task. The 32 workers in that warp cooperate, rather like a team of 32 people solving one small puzzle.
This approach is useful because CuACD often has hundreds or thousands of small geometry tasks happening at the same time.
Keeping everything on the GPU
Older approaches often had to send information back to the CPU between stages. This is similar to a factory that must stop production every time it needs to ask a manager how much storage space is available.
CuACD avoids this using a GPU memory allocator, also called a device-side heap allocator. It lets the GPU request and release memory by itself while the program is running.
As a result, the entire process—including searching for cuts, cutting meshes, building convex hulls, and measuring concavity—can stay on the GPU until the final answer is ready.
Important GPU components
The researchers create GPU versions of several geometry operations:
- Cutting a mesh with a plane
- Sorting and removing duplicate points
- Building convex hulls
- Finding connected parts of a mesh
- Measuring distances between surfaces
- Allocating temporary memory
- Searching through possible cuts
These components are also released as separate CUDA modules, so other programmers may be able to reuse them.
4. What did the researchers find?
The experiments used three groups of 3D models:
- The standard V-HACD benchmark
- Robot and household-object models from PartNet-Mobility
- A sample of models from Objaverse
The researchers compared CuACD with methods such as CoACD, NavACD, and VisACD.
Much faster processing
CuACD was reported to be dramatically faster than the comparison methods.
| Dataset | CuACD average time | CoACD average time | Approximate speedup |
|---|---|---|---|
| V-HACD | 0.23 seconds | 18.03 seconds | 78× |
| PartNet-Mobility | 0.16 seconds | 12.82 seconds | 80× |
| Objaverse subset | 0.25 seconds | 25.93 seconds | 104× |
In other words, a task that might take around 18 seconds with CoACD took about a quarter of a second with CuACD on the tested hardware.
The paper also reports that CuACD was about 40 to 64 times faster than VisACD, which was the fastest earlier GPU-assisted method included in the comparison.
Similar or better quality
Speed was not the only improvement. CuACD also produced:
- Similar or lower concavity scores
- Similar or fewer final pieces
- Cuts that often followed meaningful features, such as handles, wheels, and joints
For example, on the V-HACD benchmark:
- CoACD created about 40.5 pieces per model.
- CuACD created about 33.6 pieces.
- CuACD had a slightly lower average concavity score.
This suggests that the program was not simply making rougher results to become faster.
More detailed shapes are possible
The user can choose a stricter concavity limit. A stricter limit creates more pieces but preserves more detail.
According to the paper, lowering the threshold from 0.05 to 0.01 increased the average number of pieces from about 34 to about 194. Even so, the average processing time increased from around 0.23 seconds to about 0.5 seconds.
This is useful because users can choose between:
- Fewer, simpler pieces for faster collision checking
- More pieces for a closer representation of the original model
Strong performance on large inputs
The researchers also tested meshes with up to about one million triangles. CuACD continued to work within the memory limits of a single NVIDIA RTX 4090 GPU, although larger models naturally took longer.
5. Why are these results important?
Physics programs, video games, and robot simulators often need collision shapes before they can use 3D objects. Creating these shapes can take a long time, especially when thousands or millions of objects must be processed.
CuACD could make this preparation much faster. Possible benefits include:
- Faster loading and preparation of game assets
- Quicker robot simulation experiments
- Faster creation of collision models for large 3D datasets
- More interactive tools for artists and designers
- Less need to process models overnight on large CPU computer clusters
For example, a robotics researcher working with thousands of objects could prepare a dataset in much less time. A game designer might also be able to generate collision shapes while working instead of waiting for an offline processing step.
Conclusion
This paper presents CuACD, a GPU-based system for dividing complicated 3D models into useful convex pieces. Its main innovation is that it keeps the entire process on the GPU and organizes the work around groups of 32 GPU workers called warps.
The experiments show very large speed improvements—often more than 70 times faster than earlier CPU-based methods—while maintaining or improving the quality of the decomposition.
The wider lesson is that even complicated, irregular geometry problems can sometimes be redesigned to run efficiently on GPUs. If the method becomes widely adopted, it could make physics simulation, robotics, video games, and large-scale 3D processing much faster and more convenient.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Hardware generalizability is unresolved: the end-to-end evaluation is centered on an RTX 4090, with only a limited laptop RTX 3080 result; performance across different GPU architectures, memory capacities, CUDA versions, and non-NVIDIA accelerators is not established.
- Multi-GPU and distributed scaling are unexplored: the paper does not determine whether independent meshes, large individual meshes, or decomposition worklists can be efficiently distributed across multiple GPUs.
- CPU–GPU transfer costs are excluded from the main timing claims: the reported timings begin after preprocessing and mesh upload, so end-to-end latency from asset loading through collider export remains unknown.
- The impact of PaMO preprocessing is not fully characterized: all principal comparisons use PaMO-processed meshes and exclude its runtime, but changes in topology, thin structures, curvature, and collision-relevant details caused by preprocessing are not evaluated.
- Baseline comparisons are not fully controlled across algorithms: CuACD uses a different candidate-generation strategy and search procedure from CoACD, while the reported speedup combines GPU implementation, enriched candidate pools, and algorithmic changes; the isolated contribution of each change is not completely quantified.
- Comparisons with learning-based ACD methods are absent: RL-ACD and other learned approaches are excluded because their artifacts are unavailable, leaving the relative quality, speed, and generalization of CuACD versus current learned methods unresolved.
- Downstream physics performance is not measured: the paper reports concavity and number of parts, but does not evaluate collision-detection throughput, contact stability, penetration errors, simulation accuracy, or training performance in Bullet, PhysX, MuJoCo, or robotics simulators.
- The relationship between the selected concavity metric and practical collision quality remains uncertain: it is not shown whether lower bidirectional Hausdorff/volume-based cost consistently produces better contact behavior for diverse physical tasks.
- Hausdorff evaluation is itself approximate: distances are estimated from uniformly sampled points, but the paper does not quantify sampling error, missed worst-case deviations, or sensitivity to mesh scale and triangle density.
- The use of the cheap surrogate during tree search may alter search quality: the frequency with which surrogate-optimal cuts differ from cuts selected using the full Hausdorff-based objective is not reported.
- Quality evaluation relies primarily on aggregate means: mean concavity and mean part count can hide severe outliers; per-mesh distributions, worst-case errors, percentile runtimes, and failure-prone geometry categories are not analyzed in detail.
- The paper does not report systematic failure cases: robustness on non-manifold meshes, self-intersections, open boundaries, inverted normals, duplicate vertices, extreme coordinate ranges, very thin shells, and highly disconnected or nested components remains unclear.
- The stated watertight clipping assumptions limit applicability: boundary recovery depends on manifold edge multiplicities, but the behavior when those assumptions fail is not specified or experimentally evaluated.
- Numerical robustness of the clipping and cap triangulation stages is insufficiently characterized: nearly coplanar vertices, repeated intersections, degenerate cap polygons, narrow sliver triangles, and precision issues in plane-side tests may cause incorrect topology.
- The integer-coordinate convex-hull formulation is not fully specified: the conversion from input floating-point coordinates to integer predicates, including scaling, overflow limits, quantization error, and behavior over large coordinate ranges, is not evaluated.
- The exactness claim for convex hull construction is incomplete: although integer determinants avoid floating-point tolerance decisions, the paper does not provide formal correctness guarantees or adversarial tests for degeneracies such as coplanar, collinear, duplicated, or cospherical points.
- Allocator behavior under long and diverse workloads is unresolved: experiments use a particular stress-test distribution and a fixed 64-arena configuration, but sustained fragmentation, out-of-memory behavior, arena imbalance, allocation failures, and performance across larger candidate trees are not reported.
- The fixed reservation of 70% of free VRAM may be impractical in shared or memory-constrained environments: its effect on other GPU workloads, CUDA context coexistence, and peak-memory variability is not studied.
- Memory scaling for simultaneous large meshes is unclear: the million-triangle experiment uses a single GPU and reports aggregate behavior, but the maximum number of concurrently decomposed assets and the effect of batching heterogeneous meshes are not established.
- Warp-level load imbalance is not comprehensively quantified: occupancy is reported as 25–50%, but the paper does not separate losses from branch divergence, unequal mesh sizes, candidate-size variance, register pressure, memory latency, and allocator contention.
- The choice of warp as the universal task unit may be suboptimal for very small or very large parts: no comparison is provided against adaptive thread-, warp-, or block-level scheduling.
- The quick-sort implementation lacks broad adversarial evaluation: pivot robustness, worst-case behavior, duplicate-heavy inputs, nearly sorted inputs, and comparisons with appropriately configured radix or segmented GPU sorting are not reported.
- The direction-extreme hull prefilter may have distribution-dependent benefits: its fixed icosphere directions and threshold of 1,024 points are not justified across anisotropic, highly concave, adversarial, or nonuniform point distributions.
- Search hyperparameters may not transfer across datasets: the selected values for candidate counts, look-ahead depth, branching factor, and concave-edge sampling are evaluated on the reported benchmarks, but no principled tuning method or cross-domain robustness analysis is given.
- The threshold sweep does not establish monotonic quality guarantees: although tighter generally produces more parts, the paper does not determine whether per-part concavity, collision quality, or part count is monotonic for every input.
- The greedy hull-merge pass may be order-dependent: sensitivity to candidate ordering, merge ordering, and local minima is not analyzed, and no comparison with globally or jointly optimized merging is provided.
- The effect of connected-component splitting on semantic or physical decomposition is unknown: splitting disconnected components may be beneficial computationally, but its consequences for articulated objects, intentionally connected collision structures, and downstream asset semantics are not evaluated.
- Temporal and dynamic use cases remain unexplored: the method is evaluated for static preprocessing, while animated meshes, deforming geometry, topology changes, and incremental updates are not addressed.
- Determinism and reproducibility are not established: atomic operations, best-effort union-find updates, parallel allocation, and scheduling may produce nondeterministic outputs or runtimes, but repeatability across runs is not reported.
- Energy efficiency and total resource cost are omitted: the paper reports wall-clock speedups but does not compare GPU energy consumption, carbon cost, or total hardware cost against CPU clusters and hybrid GPU pipelines.
- Interactive authoring benefits are asserted but not measured: whether sub-second decomposition improves artist iteration, reduces manual proxy creation, or meets interactive application latency requirements remains an open empirical question.
- The reusable CUDA components are not evaluated outside CuACD: portability to other search-based ACD systems, mesh-processing pipelines, compiler configurations, and different data layouts is claimed but not demonstrated quantitatively.
Practical Applications
Immediate Applications
- Accelerated collision-geometry preprocessing for game engines (Industry—gaming and interactive 3D)
- Potential tools: a CUDA-based asset-import plugin, automatic collider generation in physics engines, or a command-line batch converter for engines such as Bullet, PhysX, MuJoCo, and game-editor toolchains.
- Workflow: upload a mesh to the GPU → select a concavity threshold and part-count target → generate convex hulls → export collider geometry and metadata.
- Dependencies: NVIDIA CUDA-compatible hardware, sufficient VRAM, and integration with the target engine’s convex-shape format. The paper’s timings exclude PaMO mesh preprocessing, so production pipelines may still need a separate simplification or remeshing step.
- High-throughput collider generation for robotic simulation (Industry/Academia—robotics and embodied AI)
- Potential products and workflows: automatic collider generation for Isaac Sim, MuJoCo, Bullet, ManiSkill, and differentiable robotics environments; dataset-preparation services that preprocess thousands of robot or household-object models.
- Practical benefit: faster construction of simulation-ready assets and more frequent regeneration of colliders when meshes or decomposition thresholds change.
- Dependencies: the generated convex parts must be validated for watertightness, contact stability, and simulator-specific limits on the number of convex bodies. GPU memory requirements may become significant for large articulated scenes or simultaneous batch processing.
- Interactive 3D content-authoring assistance (Industry—CAD, digital content creation, and game development)
- Potential tools: viewport overlays showing convex parts, automatic “generate collider” buttons, and sliders for coarse-to-fine decomposition.
- Dependencies: latency will vary with mesh complexity, GPU model, asset batching, and required output quality. The paper reports averages; worst-case latency and user-interface integration are not evaluated.
- Batch preprocessing for large 3D and robotics datasets (Academia/Industry—3D AI and simulation)
- Potential workflow: distributed GPU workers read meshes, run CuACD, export convex decompositions, and cache them alongside dataset records.
- Practical benefit: reduced CPU-cluster occupancy and faster preparation of training environments.
- Dependencies: throughput at cluster scale, GPU scheduling, data-transfer overhead, and licensing or format compatibility of source datasets. The study evaluates a single GPU extensively but does not report multi-GPU orchestration.
- Drop-in GPU acceleration for existing computational-geometry pipelines (Industry/Academia—software infrastructure)
- Potential applications: GPU mesh slicing, collision preprocessing, geometric validation, remeshing support, proximity analysis, and batched convex-hull computation.
- Potential software artifact: a CUDA library exposing kernels over structure-of-arrays mesh buffers, allowing existing GPU geometry systems to reuse the components without adopting the full CuACD search pipeline.
- Dependencies: CUDA programming expertise, careful memory-management integration, and validation against numerical and topological edge cases. Performance benefits may be smaller for workloads lacking many independent small geometry tasks.
- Faster collision-query preparation for physics-based animation (Industry—animation, virtual prototyping, and simulation)
- Potential workflow: generate multiple collider levels of detail using different values and select among them according to frame rate or simulation accuracy requirements.
- Dependencies: the relationship between the paper’s concavity metric and application-specific contact quality must be validated for each physics engine. A lower concavity does not automatically guarantee better stability for friction, joints, stacking, or thin-shell contacts.
- Teaching and research infrastructure for GPU computational geometry (Academia)
- Potential uses: laboratory exercises implementing warp-level sorting, GPU clipping, convex hulls, union-find, and device-side allocation; comparative studies of block-centric versus warp-centric designs.
- Dependencies: the code’s CUDA and GPU-generation assumptions, documentation quality, and the need to distinguish reproducible benchmark results from hardware-specific optimizations.
- Automated collision-mesh quality control (Industry—asset validation and QA)
- Potential workflow: every asset commit triggers GPU decomposition; the system reports part count, maximum concavity, memory use, and failed topology checks.
- Dependencies: project-specific acceptance thresholds and robust handling of non-manifold, self-intersecting, or incorrectly scaled meshes. The paper reports zero failures on its evaluated inputs but does not establish universal robustness.
Long-Term Applications
- Real-time or near-real-time dynamic collider generation (Long-term—games, robotics, and interactive simulation)
- Potential products: runtime destruction systems, dynamic object spawning with automatic colliders, and adaptive collision geometry for deformable robots.
- Dependencies: temporal coherence, incremental updates, bounded worst-case latency, and memory reclamation for frequently changing meshes. Recomputing a full decomposition for every frame is unlikely to be practical.
- GPU-accelerated motion planning and navigation geometry (Long-term—robotics and game AI)
- Potential workflow: generate convex parts for obstacles, combine them with navigation meshes or free-space constraints, and update the representation as environments change.
- Dependencies: the current system optimizes collision-oriented concavity rather than guaranteeing navigational correctness. Additional methods are needed to preserve clearance, avoid blocking valid paths, and support dynamic scenes.
- Physics-aware synthetic-data generation at scale (Long-term—robotics and machine learning)
- Potential outputs: training corpora containing meshes, convex decompositions, contact labels, trajectories, and rendered observations.
- Dependencies: decomposition quality must be correlated with downstream learning performance. Dataset bias, simulator fidelity, GPU availability, and robust handling of real-world scans remain open concerns.
- Learned or adaptive decomposition systems (Long-term—AI and computational geometry)
- Potential tools: learned cut proposal networks, per-object threshold predictors, or hybrid systems that use neural inference for proposal generation and exact GPU geometry for validation.
- Dependencies: training data, generalization across object categories and mesh quality, reproducibility, and maintaining geometric guarantees. The paper does not test learned proposal mechanisms.
- Multi-GPU and cloud-scale geometry services (Long-term—cloud computing and enterprise software)
- Potential architecture: a queue of mesh-decomposition jobs, GPU workers, cached multilevel collider outputs, and automatic selection of based on application requirements.
- Dependencies: batching efficiency, GPU isolation, VRAM fragmentation, fault tolerance, data-transfer costs, and support for non-NVIDIA accelerators. The current implementation is CUDA-specific and does not demonstrate distributed execution.
- Porting the warp-centric design to other irregular geometry problems (Long-term—GPU software and scientific computing)
- Potential tools: a general-purpose warp-centric geometry framework with reusable allocators, sorting, clipping, hull, and traversal abstractions.
- Dependencies: each target algorithm must expose enough independent work and tolerate SIMT control-flow divergence. Workloads with few large tasks, strong global dependencies, or highly skewed sizes may not achieve the reported gains.
- Hardware-portable GPU geometry libraries (Long-term—systems research)
- Potential outcome: portable libraries for convex decomposition and batched computational geometry in robotics, CAD, and simulation stacks.
- Dependencies: warp-level intrinsics, memory-allocation primitives, integer-predicate behavior, and performance portability differ across platforms. Porting may require redesign rather than straightforward translation.
- Formal quality and safety guarantees for physical simulation (Long-term—physics, robotics, and safety-critical engineering)
- Potential applications: certified collider generation for autonomous robots, virtual prototyping, and safety-critical digital twins.
- Dependencies: the paper demonstrates matched or improved geometric metrics and part counts, but it does not establish formal bounds on simulated dynamics. Such guarantees would require experiments across physics engines, contact models, scales, and adversarial geometries.
Glossary
- Approximate convex decomposition (ACD): The process of partitioning a nonconvex mesh into a small number of approximately convex components. “Approximate convex decomposition (ACD) converts triangle meshes into small sets of convex parts”
- AtomicCAS: An atomic compare-and-swap operation used to update shared memory safely in parallel. “we attempt the rank update with a single atomicCAS”
- Bidirectional Hausdorff distance: A geometric distance measuring the greatest deviation between two shapes in both directions. “CoACD's concavity metric unchanged. For a part with convex hull , the concavity is”
- Bitonic network: A fixed comparison network used for parallel sorting. “we apply a bitonic network to sort them directly”
- Bounding-volume hierarchy (BVH): A tree structure that groups geometric objects within nested bounding volumes to accelerate spatial queries. “the BVH scratch for Hausdorff queries”
- Concavity objective: A numerical function measuring how much a shape differs from its convex hull. “an equally expensive concavity objective---a scalar measuring how far a piece deviates from convex”
- Concave-edge plane: A cutting plane positioned using a mesh edge whose incident faces form a strongly concave angle. “Concave-edge planes target contact-relevant features”
- Connected-component analysis: The identification of separate, mutually disconnected regions in a mesh or graph. “connected-component analysis, and Hausdorff distance”
- Coalescing-binned allocator: A memory allocator that groups free blocks into size classes and merges adjacent free blocks. “The remedy is the coalescing-binned machinery”
- Computational geometry: The study and development of algorithms for geometric objects and spatial relationships. “the irregular kernels of computational geometry”
- Convex hull: The smallest convex shape containing all points of a given shape or point set. “building the convex hulls of the two resulting pieces”
- Convex collider: A convex geometric representation used for efficient collision detection in simulation. “Modern physics engines rely on convex colliders”
- CUDA: NVIDIA’s platform and programming model for general-purpose GPU computing. “We present CuACD (CUDA ACD)”
- Device-resident: Stored and processed in GPU memory without transferring intermediate data to the host CPU. “The first fully GPU-resident ACD algorithm and system”
- Dihedral angle: The angle between two planes, such as the planes of adjacent polygonal faces. “we sample up to $N_{\mathrm{ce}$ concave edges with dihedral angle ”
- Ear-clipping: A polygon-triangulation method that repeatedly removes suitable triangular “ears” from a polygon. “Bridge holes by ear-clipping into cap ”
- Floating-point tolerance: A numerical threshold used to account for rounding errors in floating-point computations. “the kernel is free of floating-point tolerance decisions”
- GPU-resident: Existing entirely in GPU memory and executing without intermediate host-CPU coordination. “a fully GPU-resident ACD algorithm”
- Heap allocator: A memory-management component that dynamically allocates and frees storage during execution. “a device-side heap allocator”
- Hybrid prefilter: A preliminary filtering step that removes points unlikely to contribute to the final convex hull before exact construction. “the hybrid variant discards them first”
- Integer determinant: A determinant computed over integer coordinates, commonly used for exact geometric orientation tests. “All predicates are integer determinants”
- Intersection-free remeshing: Mesh reconstruction that avoids introducing intersecting geometric elements. “parallel intersection-free remeshing and mesh simplification”
- Kernel: A GPU function executed concurrently by many GPU threads or warps. “the dominant costs---search, mesh cutting, and convex hull construction---have remained on the CPU”
- Lock-free: A concurrency property in which system-wide progress occurs without mutual-exclusion locks. “We implement this as a warp-parallel union-find over triangles”
- Look-ahead tree search: A bounded recursive search that evaluates possible future actions before selecting the current action. “a parallel look-ahead tree search”
- Manifold property: A geometric condition in which each local region has a topology resembling Euclidean space; for a closed triangle mesh, an interior edge is typically incident to two faces. “We use the manifold property---an interior edge is shared by exactly two triangles”
- Monte Carlo tree search (MCTS): A randomized tree-search algorithm that estimates the quality of actions through repeated simulations or sampling. “drives the search with Monte Carlo tree search”
- OptiX: NVIDIA’s GPU ray-tracing framework for accelerating ray-based geometric queries. “evaluated with OptiX ray queries on the GPU”
- Plane clipping: The operation of splitting geometry against a plane and retaining the resulting portions. “Mesh--plane clipping”
- Ray-casting: A geometric technique that shoots rays through a shape to determine containment or intersections. “Classify each into outer/inner loop by 2D ray-casting”
- SIMD: A parallel execution model in which one instruction operates on multiple data elements simultaneously. “each warp is a $32$-lane SIMD processor”
- SIMT: A GPU execution model in which groups of threads execute a common instruction stream while operating on separate data. “The irregular, data-dependent control flow of convex hull construction”
- Spatial indexing: Organizing geometric data to accelerate spatial searches and proximity queries. “rasterization, spatial indexing, hierarchy building, and traversal”
- Struct-of-arrays: A data layout that stores each field of a collection in a separate contiguous array. “CuACD's kernels instead work on plain struct-of-arrays buffers”
- Streaming multiprocessor (SM): A GPU hardware unit that schedules and executes groups of parallel threads. “a few hundred independent tasks across the GPU's streaming multiprocessors”
- Tail of stragglers: The small remainder of active work that leaves many parallel processors idle near the end of a computation. “each kernel boundary ends in a tail of stragglers”
- Temporal coherence: The exploitation of similarity between successive time steps in an animated or changing scene. “extend ACD to animated meshes via temporal coherence”
- Union-find: A data structure supporting efficient maintenance of disjoint sets and connectivity queries. “a warp-parallel union-find over triangles”
- Variable-sized buffer: A memory region whose required size is determined dynamically during execution. “the variable-size vertex arrays a plane cut emits”
- Warp: A fixed group of GPU lanes that execute instructions together, typically as a single scheduling unit. “only the warp---a fixed group of $32$ lanes sharing one instruction pointer”
- Warp-level intrinsic: A GPU instruction or primitive enabling communication or voting among lanes within a warp. “using warp-level vote and shuffle intrinsics”
- Watertight mesh: A mesh with no gaps or cracks, so its surface forms a closed boundary. “This produces two watertight sub-meshes”
- Voxelized recursive bisection: Repeatedly splitting a voxel representation of an object into smaller regions. “V-HACD replaces triangle-level clustering with a voxelized recursive bisection”
- Worklist: A dynamically maintained collection of tasks or geometric parts awaiting processing. “The algorithm maintains a worklist of parts that still require further decomposition”