WorldBench: Evaluating LLMs on Three.js Voxel World Generation
arXiv cs.GR details WorldBench, a benchmark and judge for open-ended Three.js voxel worlds generated by large language models. The core idea is simple: if a model can write an interactive 3D world, evaluation should inspect the running scene, not just a handful of renders or the source file.
That matters because the paper finds the visual and code-based views disagree on 32% of required items across five frontier models, with many misses coming from code that never appears in a frame and from small details that fixed camera views overlook. WorldBench addresses that by orbiting the world, advancing its clock, and sending a navigator agent through the scene while also checking code claims against the source and measurable pixels where possible.
The benchmark was built around a prompt for a floating voxel island with ten biomes plus physics and day/night and seasonal cycles. A mutation test, where features were removed by construction, showed code-only judging still gave full credit to four of five removed features because the code remained in the file.
For developers, the practical takeaway is that LLM-assisted worldbuilding needs evaluation that matches how players will actually experience the scene. The authors say their judge reduces credit kept on removed features from 5.44 to 3.55 out of 7.11, and they’ve released the code, prompt, tests, and judge configuration for others to inspect or extend.
“the two views disagree on 32% of required items”
- what
- WorldBench is a benchmark and judge for LLM-generated Three.js voxel worlds
- who
- Krish Bakshi; evaluated Claude Fable 5.1, GPT-6 Astra, Kimi K3, Grok 4.7, and Gemini 3.1 Pro
- when
- Submitted to arXiv on 7 Oct 2026
- impact
- Shows code-only or screenshot-only evaluation can miss important world features and overcredit broken content
Useful evaluation research, but no direct product win yet
Follow JavaScript updates
See relevant stories in your personalized news feed.
Discussion