Skip to main content
GameDev.net gamedev.net
Research Paper

This is an academic paper or technical research. Key findings may require technical background to fully understand.

Explore Research Radar

PRO Tired of ads? Read GameDev.net ad-free and help keep the community independent with GameDev Pro — $3/month.

arXiv cs.GR
arXiv cs.GR Research
· 1 month, 3 weeks ago • Yanhao Jia, Jiepeng Wang, Haibin Huang, Chi Zhang, Erik Cambria, Xuelong Li

MMOE: Modernizing Diffusion Transformers with Efficient Expert Design

Briefing

Modern large-scale AI has increasingly leaned on sparse experts to add capacity, but diffusion-transformer models used for AIGC have mostly chased bigger parameter counts and higher sparsity ratios instead of balanced efficiency. MMOE, short for ModernMOE, tries to import the practical scaling tricks that made modern LLMs workable: routed experts, shared and lightweight experts, gate-residual routing, and attention-residual reuse.

The key point for game teams is not just that the model is larger, but that it is designed to be cheaper to train and deploy relative to its quality. Every run was done on a single eight-GPU H100 node with batch size 256 for 400k steps, which is a notably accessible setup compared with the sprawling clusters often associated with foundation-model work. Under matched training and sampling settings, MMOE reached lower FID at every recorded checkpoint, meaning it converged faster than dense and intermediate sparse baselines.

Among the sparse variants, MMOE also landed the best balance between output quality and cost. The routing analysis is interesting in its own right: expert specialization stayed stable across depth, lightweight routes were used heavily, and routing changed only modestly from step to step during denoising. That suggests the model is not just throwing experts at the problem, but organizing them in a way that remains predictable during generation.

For developers watching the AI tooling space, the practical takeaway is that diffusion models may be able to follow the same “scale, but efficiently” path that LLMs took. If this...

“MMOE reaches lower FID at every recorded checkpoint”

— Paper abstract · Claim about faster convergence versus baselines
Original source
Read on arXiv cs.GR
At a glance
what
MMOE is a diffusion-transformer variant that adds routed, shared, and lightweight experts plus residual reuse to improve efficiency.
who
Yanhao Jia, Jiepeng Wang, Haibin Huang, Chi Zhang, Erik Cambria, and Xuelong Li.
when
Submitted on 27 Jul 2026; experiments trained for 400k steps.
impact
It achieved lower FID than dense and intermediate sparse baselines at every checkpoint, with a better quality-cost balance among sparse models.
Signal Positive

Promising efficiency gains without obvious quality tradeoffs

Discuss

Follow ai updates

See relevant stories in your personalized news feed.

Sign in to follow

Continue on GameDev.net

Useful next steps related to this story.

Game development news without the noise

One useful weekly briefing. No daily flood.

Sending your confirmation email…

Discussion

Loading comments...

Recommended resources

Graphics Programming Resources

See full guide
Real-Time Rendering, Fourth Edition cover
Editor pick Community pick

Real-Time Rendering, Fourth Edition

Amazon · Book

Real-Time Rendering combines fundamental principles with guidance on the latest techniques to provide a complete reference on three-dimensional interactive computer graphics. It will help you increase speed and improve image quality and learn the features and limitations of acceleration algorithms and graphics APIs. This latest fourth edition has been updated to include a chapter on virtual reality and augmented reality and covers new topics such as visual appearance, global illumination, and curves and curved surfaces. It is for anyone serious about computer graphics who wants to learn about algorithms that create synthetic images fast enough that the viewer can interact with a virtual environment.

GameDev.net may earn a commission if you purchase through these links. This helps fund the site at no extra cost to you.

Programming with wgpu in Rust cover
Editor pick

Programming with wgpu in Rust

Amazon · Book

Unlock the full power of modern graphics programming with wgpu and Rust. This comprehensive guide takes you from foundational GPU concepts to advanced real-time rendering and compute techniques—equipping you to build fast, safe, and cross-platform graphics applications. Written for intermediate to advanced Rust developers, this book provides clear explanations, hands-on examples, and detailed insights into how GPUs process and render data. You’ll explore everything from the fundamentals of buffers, shaders, and pipelines to advanced topics like deferred rendering, shadow mapping, and GPU compute workloads.

GameDev.net may earn a commission if you purchase through these links. This helps fund the site at no extra cost to you.

GameDev.net may earn a commission if you purchase through these links. This helps fund the site at no extra cost to you.