CuACD: A Fully GPU-Resident Approximate Convex Decomposition
arXiv cs.GR details CuACD, a fully GPU-resident approximate convex decomposition system aimed at the bottlenecks behind physics, collision, and robot-learning pipelines. The paper frames ACD as a standard preprocessing step that has stayed expensive because search, mesh cutting, and convex hull construction still bounce through the CPU.
CuACD’s main idea is to redesign the pipeline around warps instead of threads or thread blocks, fusing heterogeneous stages into warp-resident kernels. It also adds a device-side heap allocator so intermediate buffers can be sized and allocated on the GPU, avoiding host round trips just to prepare the next launch.
For game developers, the practical angle is obvious: faster convex decomposition means less painful collision-authoring bakes and more room for iterative workflows on complex meshes. The authors also say the reusable CUDA modules are being released open source, which could make it easier to plug GPU acceleration into existing search-based ACD pipelines.
On the V-HACD benchmark, PartNet-Mobility, and an Objaverse subset, CuACD reportedly delivers more than an order of magnitude speedup over CoACD while matching or improving quality. That puts it in the category of tooling research that could matter to physics-heavy games, simulation tools, and any pipeline that still treats convex decomposition as an overnight job.
“the first fully GPU-resident ACD system”
- what
- CuACD is a fully GPU-resident approximate convex decomposition system for triangle meshes.
- who
- Authors: Ruoxi Shi, Xinyue Wei, Fanbo Xiang, Zexiang Xu, and Hao Su.
- when
- Submitted to arXiv on 23 Sep 2026.
- impact
- Could reduce convex decomposition from tens of seconds per mesh to GPU-side workflows, easing collision and physics preprocessing.
Promising speedup for a common pipeline bottleneck
Follow graphics updates
See relevant stories in your personalized news feed.
Discussion