Paper:Pref-GRPO

Full Name:Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning

Authors:Yibin Wang et al.

arXiv:2508.20751

First submitted:2025-08-28

Code:CodeGoat24/Pref-GRPO

Core Idea:把 GRPO 中的 pointwise reward score 替换成 pairwise preference → win rate → advantage,从“追求更高分”变成“学习相对偏好”。


1. 一句话总结

Pref-GRPO 的核心不是改变 GRPO 的 policy optimization,而是改变 GRPO 使用的 reward。

传统 GRPO:

Image
  ↓
Pointwise Reward Model
  ↓
绝对分数 R(x)
  ↓
Group Normalization
  ↓
Advantage
  ↓
Policy Optimization

Pref-GRPO:

Image Group
  ↓
Pairwise Preference Reward Model
  ↓
**两两比较:A vs B
  ↓
Win Rate**
  ↓
Group Normalization
  ↓
Advantage
  ↓
Policy Optimization

核心变化:

$\boxed{\text{Pointwise Score}\rightarrow\text{Pairwise Preference}}$

也就是:

$\boxed{\text{Score Maximization}\rightarrow\text{Preference Fitting}}$

论文认为传统 pointwise reward 在 GRPO 中容易产生 Illusory Advantage(虚幻优势),进而导致 reward hacking 和训练不稳定。


2. 论文解决什么问题?

论文主要解决两个问题:

Problem 1:T2I RL 中的 Reward Hacking

传统 Flow-GRPO 会:

  1. 一个 prompt 采样 G 张图片
  2. Reward Model 给每张图片一个分数