Reward Models

How reward models turn human judgment into a score: a scalar head, Bradley–Terry training on pairwise comparisons, best-of-N reranking, reward hacking.