That’s not a typo. Four and a half months.
For context, Alibaba’s Qwen3.5 launch earlier this year already signaled China’s strategic intent in the open-weight space. But this report quantifies exactly how far the gap has closed, and the implications are uncomfortable for anyone paying premium API rates.
The Numbers That Matter
Let’s cut through the hype and look at what Mozilla actually found. The report, built on a survey of roughly 1,400 developers, OpenRouter traffic data, and third-party benchmark indices, reveals a competitive landscape that’s shifted dramatically since the July version.
The key metrics:
| Dimension | July Report | September Report | What Changed |
|---|---|---|---|
| Capability Gap | 3% performance difference | 4.4 months on task-horizon | Shifted from composite score to task-duration metric |
| Cost Comparison | GPT-4-level inference down 50x in 3 years | Kimi K3 at 30% of Fable 5’s cost | Model-to-model cost comparison |
| Usage Data | 79% developers use open models | 8 of top 10 OpenRouter models are open-weight | Platform usage rankings |
| Revenue Split | Open models ~33% of usage | Open models get 4% of revenue | The economics gap persists |
The standout: Kimi K3 from Moonshot AI scores just three points behind Anthropic’s Fable 5 on the Artificial Analysis Intelligence Index, at 30% of the cost.
Efficiency as a Weapon
Here’s the part that should genuinely worry US labs. The Great AI Reversal narrative isn’t just about export controls backfiring, it’s about what Chinese labs did when they couldn’t access Nvidia’s best hardware.
They optimized.
Brendan Burke, semiconductors analyst at Futurum Group, explained it to Fortune bluntly: “Chinese labs found algorithms that reduce the complexity of those calculations by an order of magnitude, and then achieve better results because they’re able to summarize the most relevant tokens.”
Translation: When you can’t throw more compute at a problem, you get clever about the compute you have. The US has 74% of the world’s compute according to a White House report, yet Chinese labs are matching frontier performance with a fraction of it. That’s not a hardware problem, that’s an innovation problem.
The developer community has noticed. One long-time AI user on Reddit captured the sentiment perfectly: “Qwen3.8 convinced me. The IQ3_XXS Qwen3.8 on my gaming PC performs better than the Gemini models my company was paying for.”
The “Good Enough” Threshold
The METR data reveals something fascinating about where the actual value gap sits. Closed models handle tasks taking human experts 8-12 hours. Open models handle about 7 hours currently, but that number is climbing fast.
Mozilla’s CTO Raffi Krikorian frames it as task-specific: “The premium for closed models is justified for 8-12 hour professional tasks, while open models handle routine work at dramatically lower prices.”
But here’s the uncomfortable question: How much of your actual workload needs 12-hour task capability? For most enterprises, the answer is “less than you think.”
Ameya Kanitkar, cofounder of AI measurement platform Larridin, tracked enterprise workflows and found Chinese models like GLM 5.2 and Kimi 2.6/2.7 handle 75% of engineering tasks “reasonably well” at a fifth of the cost.
DoorDash CEO Andy Fang says Kimi is “cheaper and better quality” for their needs. Cursor used Kimi to build its Composer 2 coding agent. Thomson Reuters replaced Claude with a Qwen-based model for document review.
This isn’t theoretical. The Hugging Face CEO’s assessment of China’s AI leadership looks increasingly prescient.
The Hardware Elephant in the Room
Before you rush to rip out your OpenAI subscriptions, there’s a catch, and it’s a big one.
The 4.4-month gap and the 30% cost figure are measured API-to-API at list price. When you actually try to run these models yourself, the picture changes.
The report’s hardware chart shows:
- Best open model on one server: 52.6 score (10 points below frontier)
- Best open model on one GPU: 40 score (23 points below frontier)
- Kimi K3’s native MXFP4 checkpoint: ~1.56TB across 96 shards
- vLLM’s recommendation: At least 8 GB300 GPUs minimum
Mozilla describes this as “open but not runnable by most who hold it.” Trillion-parameter models don’t run on your gaming rig. The 27B parameter models that run locally are impressive, but they’re not frontier, yet.
MiniMax M3’s open-weight release shows the trend toward more efficient architectures, and Qwen3.8 Flash Next has people genuinely excited about consumer hardware, but we’re not there yet.
The Distillation Debate
You can’t talk about Chinese AI closing the gap without addressing the elephant in the room: distillation.
The September 8 joint advisory from NSA/CISA/FBI (AA26-251A) alleges that Moonshot extracted Claude Fable 5 data to train K3 through distillation, training one model on another’s outputs. Mozilla’s report calls the claim “asserted, and unshown.” China’s foreign affairs ministry calls the accusations “groundless.”
But here’s the thing: everyone distills. As one developer pointed out, “Anthropic doesn’t generate original data either. They just hammer sites into the ground to hoover up their data.”
The more interesting question isn’t whether China distills, it’s whether the efficiency innovations would exist without the constraints. Many developers argue the export controls created the pressure that led to the algorithmic breakthroughs.
One Reddit user summarized it well: “China and elsewhere have benefited from not having access to waste money on pointless compute. They’ve focused on efficiency instead and their models might actually be economically viable.”
What the Economics Actually Say
Here’s the most uncomfortable number in the report: Open models handle about one-third of actual AI usage but capture only 4% of revenue.
The Linux Foundation reported that closed providers took 96% of model-layer revenue on OpenRouter from May-September 2025. People use open models because they’re free or cheap, but the money still flows to the closed providers.
This creates a strange dynamic. GLM-5.2’s open-weight push and ZAI’s geopolitical power play suggest Chinese labs are playing the long game, build adoption through open weights, then monetize through scale and infrastructure.
Krikorian’s framing: “We see the decision to pay for closed models as workload-specific rather than organization-specific.”
In other words, organizations aren’t choosing between open and closed. They’re choosing per task. Routine code review goes to the cheap open model. Complex architectural reasoning goes to the frontier closed model. The hybrid approach is winning.
The “Great Good-Enough” Era
One of the most telling comments from developers: “4 months ago was the ‘good enough’ point. I appreciate more capabilities and improvements, but it was usable.”
Multiple developers hit this sentiment. The “good enough” threshold passed somewhere around Qwen 3.6, and 3.8 raised the bar for what they’d accept on local hardware. The frontier labs are pushing into increasingly exotic territory, latent reasoning, trillion-parameter MoEs, autonomous agents, while most organizations just need solid code completion and document processing at reasonable prices.
Zhipu AI’s GLM-5.3 pulled off a post-training-only release that kept pace with models trained from scratch. The “zero pretraining” approach is a shot across the bow: compute efficiency is becoming the competitive moat.
What This Means for Your AI Strategy
If you’re running an organization that uses AI, the Mozilla report should trigger some honest conversations:
- Audit your workload complexity. What percentage of your AI tasks truly need frontier capability? If it’s under 50%, you’re overpaying by an order of magnitude.
- Reconsider your routing strategy. The hybrid approach, open models for routine work, closed models for complex reasoning, isn’t just cost optimization. It’s resilience. When Anthropic’s Fable 5 got restricted in June, developers who had built everything on that API learned the hard way.
- Watch the hardware curve. The 27B parameter models running on consumer GPUs are improving faster than the trillion-parameter giants. Within a year, a gaming PC might handle what required a data center today.
- Don’t dismiss the security concerns. Whether or not the distillation allegations are proven, the geopolitical reality is that Chinese models carry different risk profiles. Run them through US-based cloud providers if you’re concerned.
The KraneShares analysis puts it in investor terms: “China AI is moving from a domestic substitution story to a global price-disruption story.” Token prices will normalize. The question is whether US labs can adapt their cost structures before the equilibrium shifts.
The 4.4-month gap isn’t a comforting margin, it’s a warning shot. MiniMax M3 beating GPT-5.5 for pennies suggests the trend is accelerating, not slowing.
The next frontier isn’t raw capability. It’s cost efficiency, task-specific optimization, and the ability to deliver “good enough” at prices that make premium APIs look like luxury goods.
Mozilla’s recommendation is simple: default to open models, pay for closed models only when the task demands it. The math is hard to argue with, especially when Chinese labs keep making the open option more capable every quarter.
The gap is closing. The question is whether American AI labs have a response that doesn’t involve just throwing more compute at the problem.

