评测指标相关:
HPSv3:Towards Wide-Spectrum Human Preference Score
ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
DPO相关:
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Diffusion Model Alignment Using Direct Preference Optimization
Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models
GRPO相关:
Qwen-Image-2.0-RL Technical Report
Groupwise Reward
Pref-GRPO:Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
Reinforcing Diffusion Models by Direct Group Preference Optimization
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
TEMPFLOW-GRPO: WHEN TIMING MATTERS FOR GRPO IN FLOW MODELS
DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment
GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
混合图像+ 文字GRPO