Skip to main content
GameDev.net gamedev.net

Computer Vision News

The latest Computer Vision coverage curated for game developers.

ScribbleEdit: A Benchmark for Scribble-Only Image Editing

arXiv cs.GR details ScribbleEdit, a new benchmark for scribble-only image editing that tests whether models can infer intent from sparse user marks. The team says current VLM- and LLM-based editors …

Research arXiv cs.GR · 1 day ago
StableGrasp: Reconstructing Physically Stable Human Hand Grasps from Single Images

arXiv cs.GR details StableGrasp, a single-image hand-grasp reconstruction method that aims for physical stability, not just plausible pose. The system separates visual hand geometry from the control target, then optimizes …

Research arXiv cs.GR · 1 day ago
KASALv2: Fully Automatic 3D Rotational Symmetry Classification and Axis Localization

arXiv cs.GR details KASALv2, a fully automatic way to classify 3D rotational symmetry and localize symmetry axes without manual labels. For pose-estimation pipelines, that means cleaner symmetry priors and less …

Research arXiv cs.GR · 1 day ago
4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction

arXiv cs.GR details 4D-HOF, a feed-forward method for reconstructing hand-object interactions from coarse foundation-model estimates. It uses conditional flow matching to correct pose, rotation, and alignment errors, aiming to replace …

Research arXiv cs.GR · 2 days ago
A Simulation-Grounded Agentic VLM Framework for Wildfire Monitoring and Reporting

arXiv cs.GR details a wildfire-monitoring VLM that turns 2D simulations into labeled video episodes, then uses them as memory for training-free reporting. The system hit 51.5% exact four-tag accuracy on …

Research arXiv cs.GR · 4 days ago
4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

arXiv cs.GR details 4DCodeBench, a benchmark that tests whether agents can turn video into executable graphics code for dynamic scenes. The setup targets deformation, fluids, and fracture, and the early …

Research arXiv cs.GR · 4 days ago
Lens Flare Removal and Reconstruction

arXiv cs.GR details a lens-flare pipeline that can both remove heavy flare artifacts and reconstruct them as editable scene elements. For graphics and tools teams, the practical angle is a …

Research arXiv cs.GR · 1 week, 1 day ago
CoDimRecon: Agentic Reconstruction of Sim-Ready 3D Scenes with Deformable Curves, Surfaces, and Volumes

arXiv cs.GR details CoDimRecon, an agentic pipeline for turning multi-view RGB into sim-ready 3D scenes with rigid, articulated, and deformable assets. The system targets a long-standing gap for deformables, reconstructing …

Research arXiv cs.GR · 1 week, 2 days ago
DRHeC: Differentiable Rendering for Hand-Eye Calibration with RGB-Based Gradients

arXiv cs.GR details DRHeC, a differentiable rendering approach for hand-eye calibration that uses RGB-based gradients instead of binary masks. The method aims to improve stability and preserve internal shape cues, …

Research arXiv cs.GR · 1 week, 2 days ago
ChronoFuseGS: Multi-Temporal Gaussian Fusion with Per-Splat Persistence and Change Visualization

arXiv cs.GR details ChronoFuseGS, a multi-temporal Gaussian splatting method that merges separately trained scene models into one evolving reconstruction. For developers working on reconstruction, digital twins, or change-aware visualization, the …

Research arXiv cs.GR · 1 week, 4 days ago
Heartian: Physiology-Aware Relightable Gaussian Head Avatar

arXiv cs.GR details Heartian, a relightable Gaussian head avatar that bakes cardiac-cycle skin-color changes into facial materials. The method keeps reconstruction quality nearly unchanged while preserving recoverable rPPG signals, which …

Research arXiv cs.GR · 2 weeks ago
\phi-RIE: From Photorealistic Reconstruction to Interactive Environments

arXiv cs.GR details ϕ-RIE, a Gaussian-native pipeline that turns photorealistic 3D reconstructions into movable simulator assets. For developers, the key idea is coupling object extraction with scene completion so edited …

Research arXiv cs.GR · 2 weeks, 2 days ago
LINGO: Latent Initialization and Gradient Optimization for Sparse-view X-ray Novel View Synthesis and CT Reconstruction with 3D Gaussian Splatting

arXiv cs.GR details LINGO, a 3D Gaussian Splatting pipeline for sparse-view X-ray novel view synthesis and CT reconstruction. The method uses latent mask-space initialization and dynamic gradient scaling to cut …

Research arXiv cs.GR · 2 weeks, 3 days ago
Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene Reconstruction

arXiv cs.GR details Mira-Scene, a 3D scene reconstruction method that aligns generated objects to pixels instead of relying on sparse pose guesses. For game teams building scene tools or AI-assisted …

Research arXiv cs.GR · 2 weeks, 3 days ago
Estimating Accurate Hand Pose in Camera Space with Vision Transformer

arXiv cs.GR details a Vision Transformer hand-pose method that estimates wrist position in camera space from a single RGB image. The approach tackles depth ambiguity and pose/wrist coupling, and it …

Research arXiv cs.GR · 2 weeks, 3 days ago
NaRPA: Navigation and Rendering Pipeline for Astronautics

arXiv cs.GR details NaRPA, a ray-tracing pipeline built to generate virtual space imagery for navigation and sensor testing. The framework models passive and active vision sensors, plus a velocimeter LiDAR, …

Research arXiv cs.GR · 2 weeks, 3 days ago
Printing the Underdetermined: Materializing Multi-solutionness in Figurative Paintings

arXiv cs.GR details a pipeline that turns figurative paintings into multiple plausible 3D interpretations instead of forcing one reconstruction. It samples camera-orbit videos, rebuilds them with 3D Gaussian Splatting, then …

Research arXiv cs.GR · 3 weeks ago
SplashSplat: Reconstructing Splashing Liquids from Real-World Multi-View Videos

arXiv cs.GR details SplashSplat, a new method for reconstructing splashing liquids from real multi-view video. The team pairs a 20-scene benchmark with seven synchronized 4K cameras at 60 fps and …

Research arXiv cs.GR · 3 weeks ago
Human-aware Design Generation: Adding 3D Humans into Graphic Designs

arXiv cs.GR details a model that inserts 3D humans into partial graphic designs, aiming to make layouts feel more deliberate and visually coherent. The system predicts pose, framing, and placement …

Research arXiv cs.GR · 3 weeks, 1 day ago
CADSplat: Sparse-View 3D Gaussian Splatting Aided by CAD Models for Robust, Photorealistic Digital-Twin Reconstruction

arXiv cs.GR details CADSplat, a sparse-view 3D Gaussian splatting pipeline that uses CAD models to reconstruct photorealistic digital twins from fewer than 15 posed images. For teams building AR, inspection, …

Research arXiv cs.GR · 3 weeks, 1 day ago
Page 1 of 11 Next