Cloudflare Open-Sources Clef: Decision Models for the Agent Era
Cloudflare released Clef and Clef-flash on October 1: Apache 2.0 decision models that answer typed questions with probabilities instead of generating text. Clef-flash hits 38.8ms median latency, and the October 9 follow-up added multimodal Clef-omni plus a price cut to $0.038 per million input tokens. We break down the architecture, the benchmarks, the cost math, and when a decision model should replace an LLM call in your agent.

On October 1, Cloudflare dropped a small but signal-heavy bomb during its Birthday Week: Clef and Clef-flash — two open-source "decision models" — plus a reinforcement-learning fine-tuning platform to go with them. A week later, on October 9, Cloudflare followed up: the multimodal Clef-omni, up to 2x faster inference for Clef, and a price cut that made Clef-flash cheaper than the competition.
This is not an ordinary model launch. It is the third Big Tech entry into the "decision model" category within two weeks — after OpenAI's Decisions API and AWS's Strands Decider 2B — and Cloudflare is the only one giving away the weights. When a company whose business is selling traffic and compute starts handing out model weights for free, it's worth a serious look at what it's really after.
Not another chat model: Clef only does "multiple choice"
First, what Clef is not: it doesn't write essays, doesn't chat, doesn't generate code. It does one thing with extreme focus — you hand it a "state" (a support message, a domain, a screenshot) plus a set of typed questions: is this ticket urgent? Which team should own it? How severe is it on a scale? What comes back isn't a paragraph — it's a probability for every option.
Technically, this works through a single "prefill-only" forward pass: a frozen Qwen backbone reads the input once, then a dedicated scoring head scores all valid options in parallel. The whole step is non-autoregressive — no token-by-token generation, which means zero output tokens. In Cloudflare's words, that makes decision models "significantly faster" than autoregressive LLMs.
Three question types are supported today: yes/no ("noul"), multiple choice ("choice"), and scoring ("score"). A call looks like this (illustrative):
{
"model": "clef",
"state": "User report: checkout has been failing for every order for the last hour",
"questions": {
"urgent": { "type": "noul", "instructions": "Is this report urgent?" },
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoices, and refunds",
"technical": "Outages, errors, and configuration",
"sales": "Plans and upgrades"
}
},
"severity": {
"type": "score",
"instructions": "How severe is the customer impact?",
"criteria": ["No impact", "Minor", "Major", "Critical"]
}
}
}
What returns is a set of typed answers with probabilities, ready for your code to route, escalate, or defer to a human. Cloudflare's positioning is blunt: let agents "programmatically gather context, make decisions, and take actions" without a human sitting on every decision point.
The name is a music-theory joke: a clef at the start of a staff fixes the pitch of every note that follows, and a decision model likewise fixes the domain of context before the actions that follow. And yes, CF happens to be Cloudflare's initials too.
Third Big Tech entry into "decision models" in two weeks
Zoom out and the industrial meaning is bigger than the model itself. TypeSafe's Jev defined the category: bounded outputs, probabilities, cheap and fast, purpose-built to sit inside workflows wherever a decision is needed. Then OpenAI followed with its Decisions API, AWS shipped Strands Decider 2B, and on October 1 Cloudflare launched Clef — three giants placing their bets within about 48 hours.
But the three are playing different games. OpenAI and AWS sell API calls; Cloudflare open-sourced the weights under Apache 2.0, put them on Hugging Face, hosted them on Workers AI, and made them fully Jev-API compatible — meaning you can swap the model ID in existing Jev code and switch to Clef seamlessly.
This is strategy in the open: Cloudflare's business is the edge network and compute. Decision-model calls will inevitably run an order of magnitude more often than LLM calls (an agent may need one judgment per step), and high-frequency, low-latency inference naturally belongs on edge nodes close to users. Giving the weights away for free buys Cloudflare the "judgment layer" of the world's agents running on its infrastructure. They've done that math very carefully.
What it actually does under the hood
Per the official blog's training details, the Clef family's technical recipe is "big model for understanding, small head for deciding":
- Frozen backbone + lightweight adapters: Clef is built on Qwen3.8-27B, Clef-flash on Qwen3.5-9B, with backbone weights fully frozen — only a routing/scoring head and rank-256 low-rank adapters (LoRA) are jointly trained. Training data is internally synthesized, deliberately permuting field orders, prompts, and schema structures to force robustness against format variation.
- Two-stage attention routing: each candidate option first extracts relevant evidence from the input, then field vectors cross-attend across the full context before confidence scores are computed. A lexical prior keeps option semantics from drifting during scoring.
- Calibrated probabilities: label-smoothed cross-entropy trains classification accuracy, then Brier loss calibrates the probabilities — and this matters enormously. A decision model's outputs are consumed directly by code, so its probabilities have to mean what they say: 90% must really mean nine times out of ten, not an LLM's offhand confidence.
- RLCD as a second optimization target: Cloudflare's own Reinforcement Learning for Calibrated Decisions grants partial credit to adjacent ordinal choices, rewards fully precise record outputs, and applies a reference penalty to prevent distribution shift.
On specs, Clef offers a 64k context window (Jev's is 32k) and ships with a vision encoder — Jev only does text classification today, while Clef could read images from day one. The blog hammers this differentiator repeatedly.
Performance and pricing: let's do the math
Cloudflare published 43 evaluations at launch. The headline claim: the Clef family leads the Jev Decision Index. But vendor-published numbers deserve a discount — a mix of wins and losses is the honest shape of things. Here are the rows from the official blog's tables most worth reading:
| Benchmark | Clef 27B | Clef-flash 9B | Jev |
|---|---|---|---|
| BANKING77 (intent classification, macro-F1) | 94.20 | 90.93 | 79.74 |
| CLINC150+OOS (macro-F1) | 97.43 | 66.77 | 89.27 |
| Home appliances (case exact) | 82.95 | 97.73 | 52.27 |
| When2Call (when to call a tool) | 72.37 | 65.58 | 80.97 |
| BRIGHT (retrieval, nDCG@10) | 45.91 | 39.26 | 47.52 |
| Median latency | 209.3ms | 38.8ms | 524.1ms |
| p95 latency | 238.6ms | 122.4ms | 536.0ms |
To be fair: Clef wins classification-style benchmarks decisively (nearly 15 points ahead on BANKING77) but loses to Jev on the more reasoning-flavored When2Call and BRIGHT. That is exactly the boundary of the decision model — it is System 1 "fast thinking," not "slow thinking." Give it what it's good at; don't ask it to reason for you.
The latency numbers are the more interesting story: Clef-flash's 38.8ms median is roughly one-thirteenth of Jev's. Cloudflare's internal test on a threat-intelligence workflow is telling: given a domain, Clef fetched, rendered, and classified it end-to-end in 2.2 seconds, while the company's fastest general LLM, gpt-oss-120b, took 4.7 seconds on the same flow and returned only two classifications. Half the time, for more complete structured results.
Pricing — updated October 9 — is where this gets genuinely disruptive:
| Model | Hosted price (per million input tokens) | Notes |
|---|---|---|
| Clef-flash | $0.038 (was $0.09) | Now cheaper than Jev; hosted context cut from 64k to 24k |
| Clef | $0.24 | 64k context unchanged; up to 2x faster as of Oct 9 |
| Clef-omni | $0.15 | Multimodal, launched Oct 9 |
Two details matter. First, decision models are billed on input tokens only — there are no output tokens, because they generate no text. That's a structural change to the cost model: LLMs burn money on both ends (with output priced higher), while a decision model's bill is single-sided, and inputs can be cached and reused.
Second, the Clef-flash price cut came with a hosted context window cut from 64k to 24k. Cloudflare's data point: only 0.24% of requests exceed 24k tokens. The Hugging Face weights are untouched and still support a 256k context for self-hosters. The tradeoff is laid out with unusual transparency — long-input users either move to Clef or self-host. Open weights just became part of the pricing strategy.
Clef-omni: stuffing audio and video into one "multiple choice"
The October 9 post brought Clef-omni: a mixture-of-experts architecture on Qwen3-Omni-30B-A3B-Instruct (30B parameters, 3B active), with audio (wav/mp3), video (mp4/webm), images, and text all flowing through a single pipeline in a single call.
Before this, deciding over a video meant "pipeline hell": transcribe speech to text, split audio and visual tracks, extract frames, run separate text models, then stitch results together. Clef-omni aligns sound with frames and scores directly: text-only decisions return in about 130ms median, images in about 150ms, audio clips in a few hundred milliseconds, and a full 21-second video with sound in about 1.5 seconds.
The official example is very vibe-coding-flavored: one call carrying an installation photo, a running-sound recording, and a fan video, asking three questions — is the model/serial label visible? Does it sound like it's running smoothly? Is the fan spinning? That kind of "multi-sense quality inspection" used to need three or four models; now it's one API call.
Benchmarks hold up too: Clef-omni takes 94.8 on BANKING77 (best in the field) and 97.7 on CLINC150+OOS (also best). Though the MoE architecture scores only 63.3 on When2Call — once again confirming: don't expect complex reasoning from a decision model.
The RL fine-tuning platform: the underestimated other half
Many read the October 1 launch as "two models," but half of Cloudflare's intent sits in the RL fine-tuning platform. The logic: a decision model's greatest value isn't generic classification — it's your own business judgment. What counts as an urgent ticket, a malicious crawler, or violating content differs at every company.
The RL pipeline Cloudflare assembled uses only building blocks it already had: AI Gateway automatically turns your AI traffic into datasets → Workers AI generates rollouts against the base Clef model → Containers run RL sandboxes for scoring and replaying agent actions → a new Trainer updates the weights → the fine-tuned model redeploys to Workers AI via BYO Model. It starts as a hands-on service with the forward-deployed engineer (FDE) team, then hardens into a self-serve platform.
The ruthless part is the loop: AI Gateway is already Cloudflare's traffic front door, so dataset capture comes free. The more a customer uses it, the more accurate their private decision model gets, and the higher the switching cost. Open weights acquire users; the RL platform retains them — classic Cloudflare playbook.
Cloudflare is already dogfooding internally: its public docs repo uses Clef to detect and close spam issues automatically, the EmDash CMS moderates plugin libraries for phishing, the data-loss-prevention team scans for PII like government IDs, and the threat-intelligence team classifies domains. All of these are "high-frequency, fixed-type, rollback-safe" scenarios — squarely in a decision model's comfort zone.
Opinion: 90% of an agent's model calls don't need to "write"
For the past two years, every "judgment" we've wired into agents has been billed at "writing" prices. Decision models are the first to price the two separately.
Think about a typical agent loop: understand intent, pick which tool to call, judge whether the tool result suffices, decide the next step, check output compliance… Only one or two of those are genuinely open-ended generation; the rest is classification, routing, scoring, yes/no judgment. Yet we've been paying frontier-model prices for all of it — and tolerating non-determinism, where the same question asked twice with different wording gets different answers and downstream code has to regex its way through.
Decision models turn "judgment" into infrastructure: deterministic inputs, bounded outputs, calibrated probabilities, predictable latency. That's a paradigm-level simplification for agent engineering: the old three-piece kit of prompt engineering + output parsing + retry logic collapses into one schema-bound call whose probabilities feed straight into if-else branches.
Two cold showers, though. First, a decision model is not a smaller LLM — it's a different species. The When2Call and BRIGHT gaps say it plainly: anything needing multi-step reasoning or open exploration still belongs on a real model. Second, benchmark tables are vendor-curated. Jev wins rows in Cloudflare's own tables, no independent third party has reproduced the results yet, and you should run your own data before choosing — fortunately, open weights make that cheap.
A practical checklist for vibe coding readers
If you're building apps and agents with AI, here's how this lands in practice:
- When to switch: user-input intent classification, support-ticket routing, content-moderation scoring, an agent's internal "which tool next" selection, RAG's "is this retrieval good enough" check. The pattern: fixed options, high call frequency, latency sensitivity, rollback-safe mistakes.
- How the cost math works: a worked example — 10,000 routing decisions a day, ~2,000 input tokens each, on hosted Clef-flash: 20M tokens × $0.038/M ≈ $0.76 per day. The same volume through a general LLM (roughly $3/M input plus output tokens) starts at tens of dollars a day — two orders of magnitude apart. The money isn't even the point; the point is that 38.8ms latency makes "judge at every step" agent architectures feasible for the first time.
- When not to switch: when you need the reasoning explained ("why is this phishing"), when the options themselves are uncertain, when multi-step reasoning is required, when the output faces humans. Open-ended generation remains irreplaceable there. A healthy architecture: decision models for high-frequency judgment, big models for low-frequency deep work.
- What open weights actually mean: three things. First, self-hosting: sensitive data never leaves your network, latency is immune to public-internet jitter. Second, auditability: weights and training methods (frozen backbone + LoRA + Brier calibration) are public, not a black-box API. Third, leverage: however a vendor reprices, you always hold a version you can run yourself. Neither OpenAI's nor AWS's closed offerings give you that.
- How to start: Clef is fully Jev-API compatible — if you already use Jev, change the model ID and try it. New users can call
@cf/cloudflare/clefon Workers AI directly; OpenRouter listscloudflare/clefandcloudflare/clef-flashtoo. Take the single most painful classification step in your current agent, A/B it for a week against probability calibration and business metrics, then decide whether to expand.
Timeline and sources
- October 1, 2026 (Cloudflare Birthday Week, Thursday): the official blog published "Introducing Clef: our open-source decision models, and new RL fine-tuning platform," launching Clef (27B), Clef-flash (9B), and the RL fine-tuning platform. Dates per the Cloudflare Blog author archive.
- October 9, 2026: the official blog published "Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash," launching Clef-omni (audio/video/image/text in one pipeline), up to 2x faster Clef inference, and Clef-flash repriced to $0.038/M tokens (hosted context adjusted to 24k).
All technical parameters, benchmark figures, and prices in this article come from the two official blog posts above; benchmark results are Cloudflare's own measurements with no independent third-party reproduction yet — read them with that caveat. The competitive-landscape section (OpenAI Decisions API, AWS Strands Decider 2B, TypeSafe Jev) synthesizes public reporting.
Sources
Related articles

Meta's SWE-sweep hides 4,068 real bugs across 100 repos in 22 languages — and gives agents no issue descriptions. The best setup fixes 75.3% with a human bug report, 4.8% without. The cliff shows that finding the problem, not writing the fix, is the human moat.

On October 8 LangChain open-sourced Restock, an agent that shops in Slack and pays with Stripe's Link wallet (real $22.18 order, real refund) — while its August 18 AgentCore Payments middleware wired the x402 protocol into the LangChain toolchain for autonomous micropayments. This piece breaks down the two payment tracks, an x402 vs. Stripe decision framework, budget guardrails for one-person teams, and the security red lines.

Google, Google DeepMind, UMD and UVA propose RRSI: don't lock down what agents can change about themselves — regularize how they search. Unregularized evolution hit 92.8 evolve score but only 40.3 OOD; RRSI reached 90.5/43.6 with fewer tokens. The data-backed case that self-improvement itself needs regularization.