Skip to main content
GameDev.net gamedev.net
2 sources covering this story

3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects

A new benchmark, 3D-DefectBench, tests how well vision-language models catch fine-grained defects in generated 3D assets. Across 84 pipeline designs and about 3.2 million defect judgments, model choice mattered most, but camera setup, input type, and prompt wording also shifted results. The takeaway: automated judges need to be evaluated as full pipelines, not just as standalone models.

First reported 2 months, 1 week ago 3d generation vision-language models benchmarking automated qa
Want to discuss what this means for developers?
Open discussion
PRIMARY SOURCE
arXiv cs.GR arXiv cs.GR

3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects

2 months, 1 week ago Read source
arXiv cs.GR arXiv cs.GR 1% match

MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models

2 months, 1 week ago Read source