Skip to main content
GameDev.net gamedev.net

Multimodal News

The latest Multimodal coverage curated for game developers.

AgenticCADedit: A Stateful, Tool-Mediated Agentic Approach to Multimodal 3D CAD Editing

arXiv cs.GR details AgenticCADedit, a stateful CAD-editing pipeline that lets multimodal requests be applied as small, verifiable actions instead of regenerating a full model each time. For game teams building …

Research arXiv cs.GR · 2 days, 23 hours ago
Intervention, Not Shared Latents: Blocking Visual Shortcuts in Audio-Video Generation

arXiv cs.GR details a causal study showing audio-video generators can latch onto visual shortcuts instead of the event itself. The paper argues shared-latent designs don’t solve the problem, and that …

Research arXiv cs.GR · 5 days, 23 hours ago
When Images Look Right and Retrieve Wrong: Coverage-Guided Cross-Scale Re-Indexing for Knowledge-Faithful Generative Perception

A new multimodal indexing method, CERES, tackles a subtle failure mode in generative perception: images can look correct yet lose the concepts needed for retrieval. The system keeps scale-sensitive entities …

Research arXiv cs.GR · 1 month ago
Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models

A new benchmark, SciDraw-Bench, tries to measure whether AI can generate scientific figures that are actually usable, not just visually plausible. It focuses on 32 structured tasks across eight figure …

Research arXiv cs.GR · 2 months, 4 weeks ago
FllumaOne: A Code-Native Multimodal CAD Dataset with Executable Programs and Kernel-Validated Feature Histories

A 100,000-sample CAD dataset now ships with executable Python programs, validated feature histories, STEP output, point clouds, and renderings. For tools and technical teams, the interesting part is that each …

Research arXiv cs.GR · 3 months, 1 week ago
ANVIL: Analogies and Videos for Lecturers

ANVIL automates a workflow many teams hand-build today: turning a concept into an analogy, then into a visual screenplay, then into executable manim animation code. The paper says it can …

Research arXiv cs.GR · 4 months, 1 week ago
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

A new arXiv paper tackles a familiar multimodal headache: keeping motion, speech, and sound effects aligned in human-centric video generation. Unison splits speech from SFX in the audio stream and …

Research arXiv cs.GR · 4 months, 2 weeks ago
Listen to Rhythm, Choose Movements: Autoregressive Multimodal Dance Generation via Diffusion and Mamba with Decoupled Dance Dataset

The introduction of the LRCM framework marks a significant advancement in dance motion generation, enabling smoother and more coherent sequences. By integrating audio and text inputs, this approach enhances the …

Research arXiv cs.GR · 5 months, 3 weeks ago