Think Twice Before Quantizing That KV Cache, Here’s the Hard Data
Quantizing the KV cache on Qwen and other LLMs saves VRAM but can gut output quality. New research shows the drop is far worse than weight quantization, with creative and technical tasks taking the biggest hit.