Transformer Architecture

The Transformer, step by step: embeddings, scaled dot-product self-attention, multi-head attention, residuals and LayerNorm, the feed-forward MLP and LM head.