How DeepSeek's GRPO drops PPO's critic: sample a group of responses per prompt and use the group's mean reward as the baseline for advantages.