How LLMs stretch from 4k to 128k+ tokens: RoPE frequencies, position interpolation, NTK scaling and YaRN, efficient attention, and needle-in-a-haystack tests.