Prompt + 2 responses
Human picks the better one
Bradley-Terry model
converts scores to win probability
Loss = -log P(chosen wins)
Human raters are shown a prompt and two (or more) responses and pick the winner — sometimes with a rating scale, but pairwise comparisons are the standard because they're faster and more consistent than absolute scores. This gives a dataset of triples: prompt, winning response, losing response.