Pre-training

How an LLM learns from trillions of tokens: the next-token cross-entropy objective, one optimizer step, Chinchilla scaling laws, and what makes a run work.