Thread-Efficient Decoding for Neural Texture Compression
Neural texture compression has always looked attractive on paper: better ratios than BCn, but with a runtime cost that can make GPU decoding awkward in practice. A new approach tackles the bottleneck directly by sharing decoder MLPs across grouped textures instead of letting every texture follow its own divergent path.
The core idea is to reduce thread divergence by clustering similar textures and training a unified decoder with a gradual freezing schedule. In the reported tests, that combination cut divergence by 25% to 52% while keeping rendering quality intact. The authors also lean on CLIP embeddings to form semantic clusters, which is a notable reminder that content-aware grouping can matter as much as raw network design.
The performance numbers are the part engine teams will care about most: across more than 500 textures and multiple real scenes, the method reached up to 8.48x speedup on a Radeon RX 9070 XT versus non-shared baselines. That kind of gain could shift NTC from “interesting compression research” toward something that is easier to justify in shipping pipelines, especially where bandwidth and memory pressure are already tight.
For developers, the practical question is whether the quality/perf tradeoff holds across your own art style and hardware mix. The work suggests that decoder sharing and smarter clustering can preserve image quality while making neural compression much less expensive to run, which is exactly the kind of engineering detail that determines whether a technique stays in the lab or lands in an engine.
“reduce thread divergence by 25%-52% while preserving rendering quality”
- what
- A thread-efficient neural texture compression decoder reduces GPU divergence by sharing decoder MLPs across clustered textures.
- who
- Janarbek Matai, Sho Ikeda, Lukasz Lipski, and Takahiro Harada.
- when
- Submitted to arXiv on 28 Aug 2026.
- impact
- Reported 25% to 52% divergence reduction and up to 8.48x faster decoding on Radeon RX 9070 XT.
Promising perf gains for a known NTC bottleneck
Follow graphics updates
See relevant stories in your personalized news feed.
Discussion