方向:Pixel-space Diffusion / DiT / Text-to-Image Generation

(Papers with Code)

一句话总结:

这篇论文系统研究了如何训练大规模 Pixel-space Text-to-Image Diffusion 模型,发现直接在 pixel space 训练收敛非常慢,因此提出 Latent-to-Pixel training strategy:先在 latent space 学习生成先验,再迁移到 pixel space 做 post-training,最终实现媲美甚至超过 latent diffusion 的效果,同时获得 3.18~4.75× 推理加速。(Papers with Code)


1. 背景 Motivation

1.1 Latent Diffusion 当前主流方案

目前主流 T2I:

Image
 |
 VAE Encoder
 |
Latent z
 |
Diffusion Transformer
 |
Latent z'
 |
VAE Decoder
 |
Image

例如:

原因:

优点

latent: