The exploration–exploitation trade-off in RL: ε-greedy, optimism under uncertainty, UCB, Thompson sampling, entropy bonuses and the multi-armed bandit.