A Framework for Low-Effort Training Data Generation for Urban Semantic Segmentation
Urban semantic segmentation still suffers from the synthetic-to-real gap: even convincing renderings can miss the camera, vegetation, and architectural quirks that matter in a target domain like Cityscapes. Traditionally, closing that gap has meant more modeling, more scene dressing, and more manual work, which quickly erodes the appeal of synthetic data in the first place.
This framework takes a different route. It adapts a pre-existing diffusion model to a target domain using only imperfect pseudo-labels, then generates high-fidelity images from semantic maps produced by any synthetic dataset. In practice, that means teams can start from low-effort scenes assembled in hours, not months, and still push them toward real-domain usefulness.
The pipeline also does the unglamorous but important cleanup work: it filters weak generations, corrects image-label misalignment, and standardizes semantics across datasets. That matters for anyone building training data for games, simulation, or robotics, because the bottleneck is often not model capacity but the cost of producing enough labeled variety that actually matches the deployment domain.
Across five synthetic datasets and two real target datasets, the method reportedly improves segmentation by up to 8.0 percentage points mIoU over state-of-the-art translation approaches. The practical takeaway is straightforward: fast semantic prototyping plus generative refinement may be enough to make cheap synthetic content competitive with far more expensive handcrafted pipelines.
“low-effort sources created in hours rather than months”
- what
- A framework adapts a diffusion model to generate target-aligned urban segmentation training data from low-effort synthetic scenes.
- who
- Damjan Kalšan, Denis Zavadski, Tim Küchler, Haebom Lee, Stefan Roth, and Carsten Rother.
- when
- Submitted 13 Oct 2025; revised 27 Aug 2026; published in ICPR 2026 proceedings.
- impact
- Reported gains reach up to +8.0 percentage points mIoU over state-of-the-art translation methods.
Promising way to cut synthetic data costs while improving accuracy
Follow ai updates
See relevant stories in your personalized news feed.
Discussion