The 1M-Token Model That Fits on One GPU: DeepSeek V4 Flash Just Made Your Cloud Bill Look Silly
DeepSeek V4 Flash runs a 284B-parameter MoE model on a single AMD MI300X. Here’s how it breaks the cloud-first assumption and what it means for AI architecture.