DPO

How DPO reaches the RLHF optimum with a single supervised loss on preference pairs, no reward model or RL loop needed, plus variants IPO, KTO and ORPO.