Constrained objective
max reward, bounded KL
Lagrangian form
reward minus beta*KL
Closed-form optimum
Gibbs / softmax policy
The RLHF objective is a KL-regularized reward maximization over the whole space of possible policies :
Layer 4 · Foundations
The KL-constrained RL objective in full, why it has a closed-form optimal policy, and PPO's variance-reduction math.
Constrained objective
max reward, bounded KL
Lagrangian form
reward minus beta*KL
Closed-form optimum
Gibbs / softmax policy
The RLHF objective is a KL-regularized reward maximization over the whole space of possible policies :