Conference: CVPR 2026 arXiv: 2511.19365 Authors: Zehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang, Qi Tian Code: Official GitHub RepositoryProject: Project Page核心关键词: Pixel Diffusion / Frequency Decoupling / DiT / Pixel Decoder / Flow Matching / DCT


DeCo 的核心思想:
让 DiT 只负责低频语义建模,让一个轻量级 Pixel Decoder 负责高频细节生成,再通过 Frequency-Aware Flow Matching Loss 让模型重点学习人眼更敏感的频率成分。
整体结构可以概括成:
noisy image x_t
│
┌─────────┴─────────┐
│ │
Downsample Full Resolution
│ │
▼ ▼
DiT Pixel Decoder
│ ▲
│ │
Low-frequency Semantic
semantics guidance
│ │
└───────► AdaLN ────┘
│
▼
Pixel Velocity vθ
│
Frequency-Aware FM
│
▼
Loss
DeCo 的关键并不是简单地“增加一个 Decoder”,而是重新分配不同频率信息的建模职责:
DiT
└── Low-frequency
├── global structure
├── semantic information
└── object/layout information
Pixel Decoder
└── High-frequency
├── texture
├── edges
├── fine details
└── high-frequency signals
论文官方也明确指出,DeCo 通过让 DiT 使用下采样输入专门建模低频语义,再利用轻量 Pixel Decoder 从全分辨率输入恢复高频信息。
DeCo 直接在 pixel space 中进行生成,不依赖 VAE。
也就是说:
传统 Latent Diffusion
Image
↓
VAE Encoder
↓
Latent
↓
Diffusion Model
↓
Latent
↓
VAE Decoder
↓
Image
而 DeCo:
Image
↓
Pixel Diffusion
↓
Image