3 articles found
LG AI Research drops a 750B-parameter MoE behemoth under Apache 2.0. It crushes long-context benchmarks but pits 37B active params against models 10x smaller. Is this Korea’s sovereign AI win or a compute flex with diminishing returns?
GLM-5.2 is the third-best model overall, but its MIT license means the real magic, distillation into small, local models, hasn’t even started yet.
Xiaomi’s MiMo v2.5 hits 1000 TPS on a trillion-parameter model using commodity GPUs. Here’s the deep dive on the FP4 quantization, DFlash speculative decoding, and TileRT systems alchemy that made it possible.