1 article found
Why round-robin routing destroys LLM inference performance and how prefix-aware routing fixes it