Evaluation News
The latest Evaluation coverage curated for game developers.
A new graphics paper tackles a problem most generative pipelines still flatten into a single score: how to describe uncertainty across multiple seeds. Using 1,000 videos from 250 modern-art captions, …
A controlled study of action-conditioned world models finds that memory design matters more than replay quality suggests. With a shared video diffusion backbone and fixed action interface, the authors show …
For teams using generative AI in production, the big shift here is moving evaluation from vague “looks good” judgments to a rubric-driven system that can be scaled. QQJ calibrates LLM …
The WorldScore benchmark introduces a unified evaluation standard for world generation in games. It features 3,000 test examples across various world types, focusing on controllability, quality, and dynamics. This benchmark …