SIMD for Collision
Box3D is getting a meaningful collision-detection speedup from wide SIMD, not the usual “pack a vector into a register” style. The win comes from processing multiple edge-edge SAT tests at once, which matters once hulls get large enough that the quadratic edge-pair search becomes the bottleneck. In practice, that means complex convex shapes benefit far more than simple boxes.
The stress test here is a convex pile benchmark with 5,120 32-point hulls, nicknamed “boulders.” A box has 8 vertices and 12 edges; the boulder jumps to 32 vertices, 59 faces, and 89 edges, which explodes edge-edge combinations to 7,921 per pair. Box3D keeps SAT as its convex collision path, so it can compute separation, normals, and contact points without relying on collision margins or the GJK/EPA fallback stack.
On an AMD 7950X at 4.42 GHz, the 500-step benchmark showed full-simulation times of 40,706 ms scalar, 17,337 ms with SSE2, and 15,762 ms with AVX2-lite on one thread. At 8 threads, that was 5,292 ms scalar, 2,410 ms SSE2, and 2,277 ms AVX2-lite. Box3D currently ships SSE2 intrinsics, while the AVX2 result comes from enabling the architecture without a full 8-wide implementation.
The practical takeaway is that wide SIMD is a strong fit for collision code when the data is SoA-friendly and the workload has enough pairwise work to amortize setup costs. It barely changes box-box, but it should help destruction, debris, and other scenes that throw complex convex hulls at the narrow phase. Box3D also keeps a 128-edge cap, which helps contain the quadratic growth that makes these optimizations...
“SSE2 is over twice as fast as scalar!”
- what
- Box3D is using wide SIMD to accelerate 3D convex collision tests, especially edge-edge SAT checks.
- who
- The work is in Box3D, the 3D physics engine related to Box2D.
- when
- Benchmark results were measured over 500-step runs on an AMD 7950X at 4.42 GHz.
- impact
- SSE2 more than doubled scalar performance; AVX2-lite improved it further, mainly for complex convex hulls.
Clear performance gains for a real engine path
Follow physics updates
See relevant stories in your personalized news feed.
Discussion