1. 基本信息


2. 核心问题

语言模型可以直接计算:

$\log \pi_\theta(y\mid x),$

因此标准 DPO 可以比较 chosen 和 rejected response 的 log probability。

但扩散模型的图像概率

$p_\theta(x_0\mid c)$

难以直接计算,因为需要对所有潜在去噪路径

$x_T\rightarrow x_{T-1}\rightarrow\cdots\rightarrow x_0$

进行边缘化。论文使用扩散模型的 ELBO,将不可计算的图像 log-likelihood 转换为可计算的噪声预测误差。

核心一句话:用扩散去噪 MSE 的差值,近似语言模型中的 log-probability ratio。


3. 标准 DPO

给定 prompt (c)、preferred sample ($x_0^w$) 和 rejected sample ($x_0^l$),标准 DPO 为:

$\boxed{\mathcal L_{\mathrm{DPO}}=-\mathbb E\left[\log\sigma\left(\beta\left[\log\frac{p_\theta(x_0^w\mid c)}{p_{\mathrm{ref}}(x_0^w\mid c)}-\log\frac{p_\theta(x_0^l\mid c)}{p_{\mathrm{ref}}(x_0^l\mid c)}\right]\right)\right]}$

它要求模型相对于 reference model: