1 article found
How a Q3 quant, aggressive KV cache compression, and native MTP speculative decoding squeeze a 27B model into 16GB VRAM with real agentic coding performance.