Reasoning Model Training

Teaching models to think: chain-of-thought, RL with verifiable rewards (RLVR), the R1-style recipe, outcome vs process rewards, and self-consistency.