How RLHF made assistants like ChatGPT: start from SFT, train a reward model on human preferences, then optimize with RL under a KL penalty.