Cloudflare just dropped a quiet bomb on the AI industry, and the timing couldn’t be more calculated. While the world was still digesting the implications of decision models like Typesafe’s Jev, Cloudflare’s Workers AI team released Clef, a family of open-weight decision models that doesn’t just match the current market leader. It embarrasses it.
The numbers tell a story that’s hard to ignore: Clef leads the Jev Decision Index, clocks median latencies of 209ms (Jev: 524ms), and comes fully open-sourced under Apache 2.0. But the real revolution isn’t the benchmark scores. It’s what Clef represents, a fundamental shift in how we think about AI decision-making in production systems.
The Death of “Trust Me, It’s AI”
Open-weight models have been quietly reshaping the enterprise AI landscape for months now. Executive mentions of open models in corporate events increased sixfold year-over-year, and for good reason. AT&T already runs 40% of its AI workloads on open models, targeting 70% within a year. Digital Realty built internal chat interfaces on open models specifically because they wouldn’t feed proprietary data into frontier systems.
But here’s what’s been missing: a model category designed for decisions, not generation. LLMs are brilliant at producing text, tool calls, and creative reasoning. They’re terrible at giving you a strict, typed, probabilistic answer in under 200 milliseconds without hallucinating a justification for why it might be wrong.
Clef attacks this gap head-on. It’s a decision model, not a chatbot, not a general-purpose LLM, but a system purpose-built to classify, route, and triage with bounded structured outputs. Pass in a customer support message, and Clef returns typed answers with probabilities: 87% urgent, billing team, severity level 2. No meandering prose. No “as an AI language model” preamble. Just structured decisions your code can act on.

What Actually Makes Clef Different
The “decision model” category was already getting crowded. Jev from Typesafe, DiffusionGemma Jev, Kev-9B, Laya, each claiming to be the answer for fast, structured classification. But Clef brings three structural advantages that aren’t just incremental improvements:
Vision encoder. Jev only does text classification. Clef can take images and classify visual content. For use cases like threat intelligence, fraud detection, or content moderation, this isn’t a nice-to-have, it’s a dealbreaker for most real workflows.
64k context window. Double Jev’s 32k. When you’re processing entire support threads, code repositories, or multi-page documents for classification, context depth directly translates to decision quality.
The latency architecture. This is where Clef gets genuinely interesting. Instead of autoregressive generation, Clef uses a prefill-only pass with Qwen as the backbone, then scores valid schema choices in parallel. No token-by-token generation. The decision step is non-autoregressive, deriving schema choices directly from internal representations using a specialized two-stage attention routing process.
| Benchmark | Clef | Clef-flash | Jev |
|---|---|---|---|
| BFCL · case exact | 98.47 | 98.76 | 95.75 |
| API-Bank · accuracy | 91.93 | 93.11 | 88.19 |
| CLINC150+OOS · macro-F1 | 97.43 | 66.77 | 89.27 |
| Home appliances · case exact | 82.95 | 97.73 | 52.27 |
| Median latency · ms | 209.3 | 38.8 | 524.1 |
Clef-flash’s cluster of wins is almost more impressive than Clef’s. A 38.8ms median latency with competitive accuracy across most benchmarks means you can put it in the hot path of agentic workflows without worrying about becoming the bottleneck.
The Fine-Tuning Gambit: RL Without the PhD
Let’s address the elephant in the room. Cloudflare isn’t just releasing models, they’re launching a reinforcement learning (RL) fine-tuning platform and calling it a product. That’s either extraordinarily ambitious or slightly unhinged, depending on who you ask.
Here’s why it matters: general-purpose decision models are useful, but domain-specific decision models are transformative. Cloudflare has 15+ years of network data across trust & safety, support triage, bot detection, and threat intelligence. They’re already using Clef internally to classify website domains with Browser Run, identifying a fashion site with 95% confidence or flagging phishing attempts in 2.2 seconds, compared to 4.7 seconds for their fastest general LLM.
The RL platform leverages existing Cloudflare primitives to create a surprisingly elegant pipeline:
- AI Gateway captures your AI traffic and builds datasets automatically
- Workers AI generates rollouts against the base Clef model
- Containers provides RL sandboxes for scoring and replaying agent actions
- [NEW] Trainer updates the fine-tuned model weights
- Workers AI + BYO Model redeploys the fine-tuned model
This is the adapter-in-memory serving pattern that’s become the 2026 standard for multi-tenant AI systems. One base model, dozens of specialized adapters, hot-swapped per request. The infrastructure story here is genuinely ahead of most enterprise ML platforms.
The Open-Weights Paradox
Here’s where things get spicy. Clef is “open-source” under Apache 2.0, with weights on Hugging Face. But the open-weights movement has a definition problem. Under the strict OSI definition, a weights-only release doesn’t qualify as open-source AI, the training data, code, and methodology must also be public. Most models marketed as “open-source” are really just open-weight, and Clef is no exception.
This matters because the gap between open-weight and genuinely open AI has real architectural implications for distributed systems. When you deploy open-weight models into production, you get transparency of parameters but not necessarily transparency of behavior. You can inspect the weights, but you still can’t fully predict how the model will classify an edge case until it happens.
That said, the direction of travel is unmistakable. Meta’s Muse Spark open-weights release, Google’s Gemma 4 with improved licensing, Xiaomi’s self-trained MiMo models, the industry has collectively decided that openness is a feature, not a liability. China’s open-weight models are now just 4.4 months behind US frontier offerings, and the gap keeps narrowing.
The Training Recipe: LoRA + RLCD
Let’s get technical for a moment, because the training methodology reveals something important about where decision models are heading.
Clef freezes the Qwen backbone (27B for Clef, 9B for Clef-flash) and jointly optimizes a routing head alongside rank-256 low-rank adapters. The math here is worth understanding: for a model of that size, LoRA reduces trainable parameters to under 1% of the total, roughly 8.2 million parameters out of 1.7 billion for a comparable Qwen-scale model with rank 8. Rank 256 is more aggressive but still manageable, trading some efficiency for substantially more adaptation capacity.
The loss function combines label-smoothed cross-entropy with a Brier loss for probability calibration, a clever hybrid that penalizes both classification errors and overconfident probabilities. This matters for production decision systems where a 95% confidence on a misclassification is worse than a 70% confidence that’s correct.
The RLCD (Reinforcement Learning for Calibrated Decisions) secondary objective is the real innovation. It grants partial credit for adjacent ordinal choices (so “severe” scored as “critical” is penalized less than “minor” scored as “critical”), rewards fully precise record outputs, and applies a reference penalty to prevent distribution shift. This is closer to the GRPO-style verifiable reward approach that’s becoming standard for reasoning models than classic PPO-based RLHF, more stable, less prone to reward hacking, and reliant on deterministic scoring rather than learned reward models.
The training data itself came from synthetic datasets with permuted field orders, prompts, and schema structures, an approach that forces the model to separate semantic classification from positional biases. If you’ve ever seen a model that works great on your dev prompts but falls apart when someone reorders the fields, you understand why this matters.
What This Means for Production Systems
Here’s the pragmatic question: should you care about Clef?
The API compatibility with Jev means you can swap it in with minimal code changes. The curl request structure is identical:
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef \
-X POST \
-H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
-d '{
"model": "clef",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": {
"urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Payments, invoices, and refunds",
"technical": "Outages, errors, and configuration",
"sales": "Plans and upgrades"
}
},
"severity": {
"type": "score",
"instructions": "How severe is the customer impact?",
"criteria": ["No impact", "Minor", "Major", "Critical"]
}
}
}'
The enterprise guarantees are also worth noting: no reading, storing, or training on your requests or responses unless you opt into the fine-tuning product. That’s a meaningful differentiator when regulatory pressure and AI safety concerns are mounting across the industry.
The Real Disruption Is Structural
Let’s zoom out. The open-weights movement and its tensions aren’t just about cost savings. They’re about architectural choice. When you can run a decision model locally with 38ms latency, you stop being dependent on API calls to frontier labs for every classification task. Your system becomes more resilient, more private, and more controllable.
The MiMo-Pro social deduction benchmarks and Qwen-Drive’s autonomous driving capabilities show that open-weight models are moving beyond “good enough for simple tasks” into genuinely capable territory. Clef extends this to the decision layer, the part of AI infrastructure that actually gates business actions.
What’s particularly notable is what Clef doesn’t try to do. It doesn’t compete with frontier LLMs on open-ended reasoning. It’s not trying to be the next GPT. It’s deliberately narrow: bounded structured outputs, typed answers, calibrated probabilities. This restraint is exactly why it wins on latency and reliability. It’s a system model, not a toy.
The Bottom Line
Clef represents a bet that the future of AI infrastructure isn’t about bigger models, it’s about right-sized models that fit into production workflows with predictable performance characteristics. The open-weight approach makes this auditable and customizable, and the RL fine-tuning platform makes it adaptable to the specific decision patterns your organization actually faces.
The white-box revolution in AI decision-making is underway, and Clef is one of its most compelling artifacts yet. Whether you’re building agentic workflows, support triage systems, or threat intelligence pipelines, the ability to run transparent, fast, fine-tunable decision models is no longer a research curiosity. It’s a production reality.




