1 article found
Opus 5 scored 24% on SlopCodeBench, revealing that frontier models fail at the one thing that matters: evolving codebases over time without breaking everything.