How an LLM picks its next token: greedy vs sampling, temperature, top-k and nucleus (top-p) sampling, beam search, and repetition and length controls.