What happens after pretraining: the canonical post-training recipe of supervised fine-tuning, reward models, and preference tuning with RLHF, DPO or GRPO.