Conference: CVPR 2026 arXiv: 2511.19365 Authors: Zehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang, Qi Tian Code: Official GitHub RepositoryProject: Project Page核心关键词: Pixel Diffusion / Frequency Decoupling / DiT / Pixel Decoder / Flow Matching / DCT


image.png

image.png

1. 一句话总结

DeCo 的核心思想:

让 DiT 只负责低频语义建模,让一个轻量级 Pixel Decoder 负责高频细节生成,再通过 Frequency-Aware Flow Matching Loss 让模型重点学习人眼更敏感的频率成分。

整体结构可以概括成:

                 noisy image x_t
                       │
             ┌─────────┴─────────┐
             │                   │
       Downsample             Full Resolution
             │                   │
             ▼                   ▼
           DiT             Pixel Decoder
             │                   ▲
             │                   │
      Low-frequency          Semantic
        semantics             guidance
             │                   │
             └───────► AdaLN ────┘
                             │
                             ▼
                     Pixel Velocity vθ
                             │
                    Frequency-Aware FM
                             │
                             ▼
                          Loss

DeCo 的关键并不是简单地“增加一个 Decoder”,而是重新分配不同频率信息的建模职责:

DiT
 └── Low-frequency
      ├── global structure
      ├── semantic information
      └── object/layout information

Pixel Decoder
 └── High-frequency
      ├── texture
      ├── edges
      ├── fine details
      └── high-frequency signals

论文官方也明确指出,DeCo 通过让 DiT 使用下采样输入专门建模低频语义,再利用轻量 Pixel Decoder 从全分辨率输入恢复高频信息。


2. Motivation

2.1 Pixel Diffusion 的问题

DeCo 直接在 pixel space 中进行生成,不依赖 VAE。

也就是说:

传统 Latent Diffusion

Image
  ↓
VAE Encoder
  ↓
Latent
  ↓
Diffusion Model
  ↓
Latent
  ↓
VAE Decoder
  ↓
Image

而 DeCo:

Image
  ↓
Pixel Diffusion
  ↓
Image