Quantization & Serving

Deploying LLMs cheaply: INT8 and INT4 quantization with GPTQ and AWQ, prefill vs decode, the KV cache, continuous batching and PagedAttention in vLLM.