From USD Scenes to Knowledge Graphs: Zero-Shot Ontology Grounding with LLMs
This paper looks at a practical bottleneck in building knowledge graphs from 3D simulation scenes: matching scene objects to formal ontology classes. The authors test whether LLMs can do that grounding step in a zero-shot, training-free way for USD scenes, instead of relying on brittle hand-built dictionaries.
On a kitchen scene with 125 objects and the SOMA-HOME Ontology, the results are strong when object names are informative: 90-96% exact-match accuracy with descriptive names, and 49-89% with abbreviated names. That’s a big jump over dictionary and embedding baselines, and it suggests LLMs can be useful as a semantic layer for simulation pipelines.
The interesting part for developers is where the model is actually getting its signal. The ablation shows LLMs lean heavily on scene-graph semantics like sibling names and parent paths. If those cues are anonymized, accuracy drops to 0-6%; geometry alone only gets 4-17%. With fully opaque names, context-augmented prompting can recover up to 48%, but it’s still clearly a metadata problem more than a vision problem.
For game teams, this is less about shipping an ontology system tomorrow and more about a direction for tooling: if your USD or scene graph data is richly labeled, LLMs may help automate classification, validation, or content indexing. If your asset naming is messy, the paper is a reminder that no amount of model cleverness fully replaces consistent authoring conventions.
“LLMs achieve 90-96% exact-match accuracy with descriptive names.”
- what
- LLMs were tested as a zero-shot, training-free way to ground USD scene objects to ontology classes for knowledge graph construction.
- who
- Jiangtao Shuai, Zongxiong Chen, Manfred Hauswirth, and Sonja Schimmler.
- when
- Submitted to arXiv on 8 Jun 2026 (arXiv:2606.09134).
- impact
- Could reduce manual dictionary work in simulation/content pipelines, especially where scene metadata is already structured and descriptive.
Strong results, but only when metadata is already good.
Follow knowledge-graph updates
See relevant stories in your personalized news feed.
Discussion