评测指标相关:

Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

HPSv3:Towards Wide-Spectrum Human Preference Score

ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding

UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture

DPO相关:

Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Diffusion Model Alignment Using Direct Preference Optimization

Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models

GRPO相关:

Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement Learning

Qwen-Image-2.0-RL Technical Report

Groupwise Reward

Pref-GRPO:Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning

Reinforcing Diffusion Models by Direct Group Preference Optimization

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

TEMPFLOW-GRPO: WHEN TIMING MATTERS FOR GRPO IN FLOW MODELS

DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment

GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping

混合图像+ 文字GRPO