How policy-gradient methods improve a policy directly: the policy gradient theorem, the log-derivative trick, REINFORCE, baselines, advantages and GAE.