Skip to main content
GameDev.net gamedev.net

Vision Language Models News

The latest Vision Language Models coverage curated for game developers.

A Simulation-Grounded Agentic VLM Framework for Wildfire Monitoring and Reporting

arXiv cs.GR details a wildfire-monitoring VLM that turns 2D simulations into labeled video episodes, then uses them as memory for training-free reporting. The system hit 51.5% exact four-tag accuracy on …

Research arXiv cs.GR · 3 days, 9 hours ago
FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos

A new vision-language system can turn origami demonstration videos into executable folding programs, using geometry checks and physical simulation to keep the sequence plausible. For game teams, the interesting part …

Research arXiv cs.GR · 1 month ago
Content Based Video Narration of Gameplay with Vision Language Models

A new VLM-driven pipeline can generate esports-style commentary from raw gameplay video without engine hooks, telemetry, or game-specific training. It packs nine frames into one image, reuses recent narration as …

Research arXiv cs.GR · 1 month, 3 weeks ago
ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning

ThinkBLOX is a new VLM-driven pipeline for generating 3D indoor scenes through progressive reasoning instead of one-shot layout planning. It iteratively places and refines objects, aiming to reduce the awkward …

Research arXiv cs.GR · 2 months, 3 weeks ago
StructuredEdit: Constraint-Aware Graphic Design Editing via Differentiable Parameter Propagation

StructuredEdit reframes graphic design editing as parameter manipulation instead of pixel generation, and that shift matters for teams shipping layout-heavy tools. The system uses differentiable parameter propagation to enforce hard …

Research arXiv cs.GR · 3 months ago
Function2Scene: 3D Indoor Scene Layout from Functional Specifications

A new paper tackles indoor scene generation from the angle designers actually care about: whether a room supports the people using it. Function2Scene turns natural-language briefs into layout constraints, then …

Research arXiv cs.GR · 4 months ago
3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code

A new benchmark, 3DCodeBench, tests whether VLM agents can turn text and image prompts into procedural 3D code that actually works in modeling software. For technical artists and tools programmers, …

Research arXiv cs.GR · 4 months ago
HOLODECK 2.0: Vision-Language-Guided 3D World Generation with Editing

HOLODECK 2.0 pushes AI-assisted 3D world building toward practical game production, combining vision-language parsing with generative asset creation and layout editing. It targets both indoor and open-world scenes, with styles …

Research arXiv cs.GR · 9 months, 2 weeks ago