Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation
UniCaMo is a unified video-generation framework aimed at a problem that matters to anyone building AI-assisted content tools: precise control over both subject motion and camera movement in a single pass. Instead of bolting on extra conditioning networks, it works directly on the diffusion model’s initial noise tensor, which is where the generation process begins.
The core trick is to build a shared 3D-grounded noise space across frames. Sparse 3D point tracks warp the reference frame’s Gaussian noise along intended object trajectories, while a virtual spherical noise representation fills in newly exposed regions as the camera moves. That combination is meant to preserve both temporal continuity and geometric consistency, even when the viewpoint changes and the scene reveals new content.
For developers, the practical appeal is that UniCaMo requires no architectural changes to the base video model. It relies on lightweight LoRA fine-tuning on large pretrained systems, including Wan 2.1 at 14B parameters, which makes it easier to slot into existing diffusion pipelines than methods that need custom adapters or control branches.
The team reports state-of-the-art results on controllable video generation benchmarks for both visual quality and motion controllability. For game teams, this kind of technique could eventually matter for previs, cinematic prototyping, synthetic animation, and rapid iteration on camera language—especially where consistent motion across generated shots is the bottleneck.
“requires no auxiliary adapters, control branches, or architectural changes”
- what
- UniCaMo is a controllable video generation framework that directly shapes diffusion input noise to control object and camera motion together.
- who
- Long Vu, Tan Ngo, Animesh Karnewar, Amir Habibian, Binh-Son Hua, Hung Bui, Minh Hoai Nguyen, and Phong Nguyen-Ha.
- when
- Submitted to arXiv on 2 July 2026.
- impact
- It may simplify AI video workflows for previs, cinematic prototyping, and motion-consistent content generation without changing model architecture.
Promising control gains with minimal model changes
Follow AI updates
See relevant stories in your personalized news feed.
Discussion