Real-Time AttentionBender: Granular Interactive Network Bending of Video Diffusion Transformers
Real-Time AttentionBender is a research prototype for live, granular control over video diffusion transformers. Instead of treating the model as a prompt-only black box, it exposes self-attention, cross-attention, and feed-forward components as separate surfaces that can be manipulated during generation.
The paper says the tool reaches down to individual diffusion steps, DiT layers, prompt tokens, and hidden neurons. That level of access matters because it gives creators a way to see which parts of the network are shaping the output, and to steer those parts directly while the video is being generated.
The authors frame this as both an XAI-for-arts probe and an expressive instrument. In practical terms, that means it could be useful for technical artists and graphics programmers who want more than prompt engineering: it offers a way to explore how model internals map to visual style, motion, and failure modes.
It is built as a plugin in the DayDream Scope ecosystem and wraps open-source real-time Wan pipelines. The paper was submitted on 24 Apr 2026 and revised on 8 Jun 2026. Even if this never ships as a production tool, it points toward a workflow where generative video systems become interactive creative tools rather than opaque output engines.
“the model's material process”
- what
- Real-Time AttentionBender is a live tool for granular interactive control of video diffusion transformers.
- who
- Authors: Adam Cole, Rebecca Fiebrink, and Mick Grierson.
- when
- Submitted 24 Apr 2026; revised 8 Jun 2026 (v2).
- impact
- Could help technical artists and graphics programmers inspect and steer generative video models beyond prompt-only control.
Promising creative control and model transparency for artists.
Follow generative-video updates
See relevant stories in your personalized news feed.
Discussion