The Markov decision process behind RL: states, actions, transitions, rewards and discounting, the Markov property, the Bellman equation and value iteration.