This is the same closed-form result from the RLHF star: for any fixed reward , the KL-constrained optimum is a reweighting of the reference by exponentiated reward. DPO's derivation inverts this relationship — treating it as defining as a function of and :
Layer 4 · Foundations
Direct Preference Optimization
The full derivation from the KL-constrained RLHF objective to the DPO loss, and its gradient's implicit reward-model interpretation.
15 min read180 XP