Vision Language Models News
The latest Vision Language Models coverage curated for game developers.
arXiv cs.GR details a wildfire-monitoring VLM that turns 2D simulations into labeled video episodes, then uses them as memory for training-free reporting. The system hit 51.5% exact four-tag accuracy on …
A new vision-language system can turn origami demonstration videos into executable folding programs, using geometry checks and physical simulation to keep the sequence plausible. For game teams, the interesting part …
A new VLM-driven pipeline can generate esports-style commentary from raw gameplay video without engine hooks, telemetry, or game-specific training. It packs nine frames into one image, reuses recent narration as …
ThinkBLOX is a new VLM-driven pipeline for generating 3D indoor scenes through progressive reasoning instead of one-shot layout planning. It iteratively places and refines objects, aiming to reduce the awkward …
StructuredEdit reframes graphic design editing as parameter manipulation instead of pixel generation, and that shift matters for teams shipping layout-heavy tools. The system uses differentiable parameter propagation to enforce hard …
A new paper tackles indoor scene generation from the angle designers actually care about: whether a room supports the people using it. Function2Scene turns natural-language briefs into layout constraints, then …
A new benchmark, 3DCodeBench, tests whether VLM agents can turn text and image prompts into procedural 3D code that actually works in modeling software. For technical artists and tools programmers, …
HOLODECK 2.0 pushes AI-assisted 3D world building toward practical game production, combining vision-language parsing with generative asset creation and layout editing. It targets both indoor and open-world scenes, with styles …