Prompt + chosen + rejected
Score both under policy
Score both under frozen reference
Compare the two log-ratios
Push chosen's advantage up
Recall from the RLHF derivation that the optimal KL-constrained policy has the form . Solving this for gives an *implicit reward* defined entirely by how much more likely the policy makes a response than the reference does: