Computer Vision News
The latest Computer Vision coverage curated for game developers.
arXiv cs.GR details ScribbleEdit, a new benchmark for scribble-only image editing that tests whether models can infer intent from sparse user marks. The team says current VLM- and LLM-based editors …
arXiv cs.GR details StableGrasp, a single-image hand-grasp reconstruction method that aims for physical stability, not just plausible pose. The system separates visual hand geometry from the control target, then optimizes …
arXiv cs.GR details KASALv2, a fully automatic way to classify 3D rotational symmetry and localize symmetry axes without manual labels. For pose-estimation pipelines, that means cleaner symmetry priors and less …
arXiv cs.GR details 4D-HOF, a feed-forward method for reconstructing hand-object interactions from coarse foundation-model estimates. It uses conditional flow matching to correct pose, rotation, and alignment errors, aiming to replace …
arXiv cs.GR details a wildfire-monitoring VLM that turns 2D simulations into labeled video episodes, then uses them as memory for training-free reporting. The system hit 51.5% exact four-tag accuracy on …
arXiv cs.GR details 4DCodeBench, a benchmark that tests whether agents can turn video into executable graphics code for dynamic scenes. The setup targets deformation, fluids, and fracture, and the early …
arXiv cs.GR details a lens-flare pipeline that can both remove heavy flare artifacts and reconstruct them as editable scene elements. For graphics and tools teams, the practical angle is a …
arXiv cs.GR details CoDimRecon, an agentic pipeline for turning multi-view RGB into sim-ready 3D scenes with rigid, articulated, and deformable assets. The system targets a long-standing gap for deformables, reconstructing …
arXiv cs.GR details DRHeC, a differentiable rendering approach for hand-eye calibration that uses RGB-based gradients instead of binary masks. The method aims to improve stability and preserve internal shape cues, …
arXiv cs.GR details ChronoFuseGS, a multi-temporal Gaussian splatting method that merges separately trained scene models into one evolving reconstruction. For developers working on reconstruction, digital twins, or change-aware visualization, the …
arXiv cs.GR details Heartian, a relightable Gaussian head avatar that bakes cardiac-cycle skin-color changes into facial materials. The method keeps reconstruction quality nearly unchanged while preserving recoverable rPPG signals, which …
arXiv cs.GR details ϕ-RIE, a Gaussian-native pipeline that turns photorealistic 3D reconstructions into movable simulator assets. For developers, the key idea is coupling object extraction with scene completion so edited …
arXiv cs.GR details LINGO, a 3D Gaussian Splatting pipeline for sparse-view X-ray novel view synthesis and CT reconstruction. The method uses latent mask-space initialization and dynamic gradient scaling to cut …
arXiv cs.GR details Mira-Scene, a 3D scene reconstruction method that aligns generated objects to pixels instead of relying on sparse pose guesses. For game teams building scene tools or AI-assisted …
arXiv cs.GR details a Vision Transformer hand-pose method that estimates wrist position in camera space from a single RGB image. The approach tackles depth ambiguity and pose/wrist coupling, and it …
arXiv cs.GR details NaRPA, a ray-tracing pipeline built to generate virtual space imagery for navigation and sensor testing. The framework models passive and active vision sensors, plus a velocimeter LiDAR, …
arXiv cs.GR details a pipeline that turns figurative paintings into multiple plausible 3D interpretations instead of forcing one reconstruction. It samples camera-orbit videos, rebuilds them with 3D Gaussian Splatting, then …
arXiv cs.GR details SplashSplat, a new method for reconstructing splashing liquids from real multi-view video. The team pairs a 20-scene benchmark with seven synchronized 4K cameras at 60 fps and …
arXiv cs.GR details a model that inserts 3D humans into partial graphic designs, aiming to make layouts feel more deliberate and visually coherent. The system predicts pose, framing, and placement …
arXiv cs.GR details CADSplat, a sparse-view 3D Gaussian splatting pipeline that uses CAD models to reconstruct photorealistic digital twins from fewer than 15 posed images. For teams building AR, inspection, …