Global Pass Barriers Without Per-Resource RHI Tracking: A Cross-Vendor Study with Blade
Blade is being positioned as a lighter-weight alternative to state-tracking RHIs like wgpu: Vulkan images stay in GENERAL, per-resource state is not tracked, and synchronization is handled with global barriers at pass boundaries. The study isolates barrier placement and stage/access scope, then compares matched wgpu programs across six GPUs from four vendors, with exploratory Apple/Metal results in the mix.
The biggest wins show up when redundant barriers are removed from independent compute passes: GPU span drops 29.3% on an RTX 5070 and 32.3% on an RX 7900 XT. Independent render workloads also improve on those discrete parts, though by smaller amounts, while a Radeon 780M case gets worse by 42.4% at 32 passes. That makes the result useful, but very workload- and stability-floor-dependent.
A second finding is more subtle: deriving global barrier scope from neighboring pass kinds, without tracking any resource, still saves 5.0% on an NVIDIA graphics chain and 6.7% on an AMD compute chain. No AMD render-involving scope cell met the stability criterion, and dependent-chain placement effects never cleared it either. In other words, the cheap aggregate approach can work, but only in some cells of the design space.
For engine developers, the practical implication is that command count alone is a poor predictor of synchronization cost. Driver behavior can expand broad dependencies into multiple flush and invalidate requests, and some of those may be elided internally. Blade’s direction is a hybrid one: keep the RHI lightweight and pass-kind aware, while leaving aliasing,...
“Removing fifteen redundant barriers from sixteen independent compute passes reduces GPU span by 29.3%.”
- what
- A study tested Blade, a tracking-free RHI that uses global pass-boundary barriers instead of per-resource state tracking.
- who
- Dzmitry Malyshau compared Blade against matched wgpu programs across six GPUs from four vendors.
- when
- The paper was submitted on 29 Jul 2026 as arXiv:2607.26506.
- impact
- Removing 15 redundant barriers from 16 independent compute passes cut GPU span by 29.3% on RTX 5070 and 32.3% on RX 7900 XT.
Promising wins, but some workloads regressed badly.
Follow graphics updates
See relevant stories in your personalized news feed.
Discussion