Taming Real-World Space-Time Video Super-Resolution with One-Step Diffusion
The main change here is moving space-time video super-resolution into a one-step diffusion setup, instead of relying on heavier multi-step generation. OSDEnhancer is designed for real-world degradations, not just clean synthetic downsampling, so it targets the kind of blur, noise, compression, and inconsistency that show up in production footage.
For developers, the interesting part is the architecture split: a linear initialization to establish structure, then two LoRAs that specialize in temporal coherence and texture enrichment, plus a bidirectional VAE decoder with deformable recurrent blocks for better latent-to-pixel reconstruction. The paper was submitted on 28 Jan 2026 and revised on 19 May 2026. If the claims hold up, this could be useful anywhere you need higher-res video with stable motion, especially for cinematic pipelines and AI-assisted content workflows.
“the first framework that achieves robust STVSR in one-step diffusion”
- what
- OSDEnhancer is a one-step diffusion framework for real-world space-time video super-resolution.
- who
- Authors: Shuoyan Wei, Feng Li, Chen Zhou, Runmin Cong, Yao Zhao, and Huihui Bai.
- when
- Submitted 28 Jan 2026; revised 19 May 2026.
- impact
- Could improve upscaling of game video, cutscenes, trailers, and captured footage while preserving temporal coherence.
Promising quality gains for real-world video upscaling
Follow video-super-resolution updates
See relevant stories in your personalized news feed.
Discussion