长久以来, 人们认为视觉Encoder重建质量越高,生成效果会越好, 该论文引入iFID:插值FID——一种与扩散模型生成质量强相关的新指标, 揭示了连通潜在空间对生成的意义

UniTok: A Unified Tokenizer for Visual Generation and Understanding

[RAE: Diffusion Transformers with Representation Autoencoders](https://aboard-hurricane-478.notion.site/RAE-Diffusion-Transformers-with-Representation-Autoencoders-3401baf34fa5812b87eed4e03068b60a)

Beyond Language Modeling:An Exploration of Multimodal Pretraining

SVG:Latent Diffusion Model without Variational Autoencoder

The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding

PS-VAE: Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing

Making Reconstruction FID Predictive of Diffusion Generation FID:

Let ViT Speak: Generative Language-Image Pre-training