How RLVR trains reasoning models with rewards a program can check: verifiers, R1-Zero's recipe, reward hacking, DAPO, the pass@k debate, and rubric rewards.