Paper: Qwen-Image-2.0-RL Technical Report
Authors: Yixian Xu et al., Qwen Team
arXiv: 2606.27608
Date: 2026.06
Task: Text-to-Image Generation + Image Editing
核心关键词: RLHF / Flow-GRPO / Reward Model / Composite Reward / Hybrid CFG / Reward Hacking / Prompt Curation / On-Policy Distillation / Velocity Matching
Qwen-Image-2.0-RL 的核心思想:
不直接把 T2I 和 Image Editing 混在一起做 RL,而是分别训练两个 task-specific RL teacher,然后通过 On-Policy Distillation(OPD)把两个 teacher 的能力蒸馏回一个统一模型。
整体流程:
Qwen-Image-2.0 Base
│
├──────────────────────┐
│ │
▼ ▼
T2I RL Teacher Editing RL Teacher
│ │
│ │
Alignment Reward Instruction Reward
Aesthetic Reward Face Identity Reward
Portrait Reward
│ │
└──────────┬───────────┘
▼
On-Policy Distillation
│
▼
Qwen-Image-2.0-RL
论文最终得到:
其中 T2I Elo +78,Editing Elo +93。
Qwen-Image-2.0 本身已经通过 supervised pretraining 获得了很强的生成能力。
但是:
Supervised diffusion training ≠ human preference optimization
传统 Flow Matching / Diffusion Training 主要优化:
预测正确 velocity
↓
MSE / Flow Matching Loss