Why BERT, GPT and T5 behave so differently: encoder-only, decoder-only and encoder–decoder Transformers, cross-attention, teacher forcing and exposure bias.