How LLMs vary attention: causal and sliding-window masks, multi-head vs multi-query vs grouped-query (GQA) and MLA, sparse attention and FlashAttention.