Faster LLM inference with identical output: a small draft model proposes tokens, the large model verifies them in one forward pass, plus Medusa and EAGLE.