arXiv:2607.17585
Authors: Renye Yan et al.
Topic: Pixel-space Diffusion + Transformer + Unified Multimodal Generation
当前主流图像生成模型(Stable Diffusion、DiT、FLUX、SD3 等)主要采用:
Image
↓
VAE Encoder
↓
Latent Space
↓
Diffusion Transformer
↓
VAE Decoder
↓
Image
即:
Latent Diffusion Model (LDM)
优点:
但是存在问题:
VAE 是固定视觉 tokenizer:
pixel space
↓
VAE
↓
latent representation
压缩过程中: