Skip to main content
GameDev.net gamedev.net

Text To Image News

The latest Text To Image coverage curated for game developers.

NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts

A new study shows that fine-tuning open-source diffusion models on 1,000 captioned nuclear-energy images can materially improve text-to-image accuracy for specialized technical prompts. SDXL benefited the most, while SD-v3.5-Medium saw …

Research arXiv cs.GR · 1 month, 2 weeks ago
Latent-Identity Tuning in Text-to-Image Personalization Models

A new latent-identity tuning method aims to make text-to-image personalization far more precise for facial edits. It works inside a frozen encoder, finding semantic directions in latent tokens instead of …

Research arXiv cs.GR · 2 months, 1 week ago
Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models

A new benchmark, SciDraw-Bench, tries to measure whether AI can generate scientific figures that are actually usable, not just visually plausible. It focuses on 32 structured tasks across eight figure …

Research arXiv cs.GR · 2 months, 3 weeks ago
Token-to-Token Alignment of Text Embeddings for Semantic Blending

A new paper argues that prompt embeddings for text-to-image models already contain a usable semantic continuum, but the tokens are misaligned. By first rephrasing prompts into a shared structure and …

Research arXiv cs.GR · 3 months ago
Bokeh Diffusion: Defocus Blur Control in Text-to-Image Diffusion Models

For graphics teams, the interesting bit is explicit control over defocus blur in text-to-image diffusion instead of hoping prompt wording fakes depth of field. The paper’s Bokeh Diffusion keeps scene …

Research arXiv cs.GR · 3 months, 2 weeks ago
Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive Smoothing

A new arXiv paper tackles a practical problem for teams using text-to-image diffusion models: proving ownership even when a suspect model has had watermark signals damaged or stripped. Cert-LAS uses …

Research arXiv cs.GR · 3 months, 4 weeks ago
On-the-fly Repulsion in the Contextual Space for Rich Diversity in Diffusion Transformers

A new diffusion-transformer sampling trick aims to fix the “samey output” problem without the usual quality hit. By pushing repulsion into the model’s contextual attention space during the forward pass, …

Research arXiv cs.GR · 5 months, 3 weeks ago
LGTM: Training-Free Light-Guided Text-to-Image Diffusion Model via Initial Noise Manipulation

A new approach to text-to-image generation, LGTM, allows developers to manipulate initial noise for better lighting control without extensive training. This innovation could significantly streamline workflows for graphics programmers and …

Research arXiv cs.GR · 6 months ago
Diverse Text-to-Image Generation via Contrastive Noise Optimization

The introduction of Contrastive Noise Optimization marks a significant advancement in text-to-image generation, addressing the common issue of limited diversity in outputs. This method reshapes initial noise rather than optimizing …

Research arXiv cs.GR · 6 months, 1 week ago
Paris: A Decentralized Trained Open-Weight Diffusion Model

A new open-weight diffusion model called Paris has been trained entirely through decentralized computation, without synchronized gradients or a central GPU cluster. It uses eight expert models and a lightweight …

Research arXiv cs.GR · 8 months, 1 week ago