1 article found
UkisAI’s post-trained Qwen cuts thinking tokens by 58% and doubles speed with under 1% accuracy loss. Here’s how they did it without crippling the model.