ExLlamaV3 v1.0.0 Just Rewrote the Rules for Local LLM Inference, NVIDIA GPU Owners Rejoice
The first stable release of ExLlamaV3 brings decode speed improvements up to 109% on RTX 5090, drops flash-attention-2 dependencies, and adds a new attention kernel with online cache quantization. Here’s what actually changed.