Qwen-Image-2.0-RL Technical Report

Paper: Qwen-Image-2.0-RL Technical Report

Authors: Yixian Xu et al., Qwen Team

arXiv: 2606.27608

Date: 2026.06

Task: Text-to-Image Generation + Image Editing

核心关键词: RLHF / Flow-GRPO / Reward Model / Composite Reward / Hybrid CFG / Reward Hacking / Prompt Curation / On-Policy Distillation / Velocity Matching


1. 一句话总结

Qwen-Image-2.0-RL 的核心思想:

不直接把 T2I 和 Image Editing 混在一起做 RL,而是分别训练两个 task-specific RL teacher,然后通过 On-Policy Distillation(OPD)把两个 teacher 的能力蒸馏回一个统一模型。

整体流程:

Qwen-Image-2.0 Base
        │
        ├──────────────────────┐
        │                      │
        ▼                      ▼
   T2I RL Teacher         Editing RL Teacher
        │                      │
        │                      │
 Alignment Reward         Instruction Reward
 Aesthetic Reward         Face Identity Reward
 Portrait Reward
        │                      │
        └──────────┬───────────┘
                   ▼
          On-Policy Distillation
                   │
                   ▼
          Qwen-Image-2.0-RL

论文最终得到:

其中 T2I Elo +78,Editing Elo +93。


2. 这篇论文到底解决什么问题?

Qwen-Image-2.0 本身已经通过 supervised pretraining 获得了很强的生成能力。

但是:

Supervised diffusion training ≠ human preference optimization

传统 Flow Matching / Diffusion Training 主要优化:

预测正确 velocity
        ↓
MSE / Flow Matching Loss