Paper:Pref-GRPO
Full Name:Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
Authors:Yibin Wang et al.
arXiv:2508.20751
First submitted:2025-08-28
Code:CodeGoat24/Pref-GRPO
Core Idea:把 GRPO 中的 pointwise reward score 替换成 pairwise preference → win rate → advantage,从“追求更高分”变成“学习相对偏好”。
Pref-GRPO 的核心不是改变 GRPO 的 policy optimization,而是改变 GRPO 使用的 reward。
传统 GRPO:
Image
↓
Pointwise Reward Model
↓
绝对分数 R(x)
↓
Group Normalization
↓
Advantage
↓
Policy Optimization
Pref-GRPO:
Image Group
↓
Pairwise Preference Reward Model
↓
**两两比较:A vs B
↓
Win Rate**
↓
Group Normalization
↓
Advantage
↓
Policy Optimization
核心变化:
$\boxed{\text{Pointwise Score}\rightarrow\text{Pairwise Preference}}$
也就是:
$\boxed{\text{Score Maximization}\rightarrow\text{Preference Fitting}}$
论文认为传统 pointwise reward 在 GRPO 中容易产生 Illusory Advantage(虚幻优势),进而导致 reward hacking 和训练不稳定。
论文主要解决两个问题:
传统 Flow-GRPO 会: