Nearly every AI model you have met answers in sentences. You ask, it writes. That is the right shape for a conversation and the wrong shape for a great many jobs a computer actually needs doing.
Cloudflare Clef is a different kind of model. Ask it a question and it does not compose a reply: it picks an answer from a list you gave it, and tells you how confident it is. Cloudflare calls this a decision model, and it is worth understanding even if you never use Clef itself, because the category is going to matter.
Checked on 10 October 2026 against Cloudflare's announcement, published 1 October 2026. Every figure here is Cloudflare's own, measured on its own benchmarks.
What is a decision model?
Start with the problem. A support ticket arrives and something has to decide which team it goes to. There are eleven teams. The answer is one of eleven things, and it is never a paragraph.
You can do this with an ordinary language model, and plenty of people do. You write a prompt listing the teams, ask for the team name back, and then you write code to handle all the ways the model fails to cooperate: it returns a team that does not exist, or wraps the name in an explanation, or formats it differently today than yesterday. The model is generating free text and you are trying to squeeze it back into a fixed shape.
A decision model removes that step. The set of allowed answers is part of the request, the model's output is one of them by construction, and it comes with a probability attached. Cloudflare describes the output as "typed answers with probabilities".
The probability is the part people overlook, and it is what makes the whole thing useful for agents. If the model is 98% sure, act on it. If it is 51% sure, send it to a human. You cannot build that rule on top of a model that always sounds equally confident.
Clef and Clef-flash
Cloudflare released two sizes, both under the Apache 2.0 licence, so you can download and run them yourself.
| Clef | Clef-flash | |
|---|---|---|
| Base model | Qwen 3.8-27B, frozen backbone | Qwen 3.5-9B, frozen backbone |
| Median latency | 209.3 ms | 38.8 ms |
| Context window | 64,000 tokens | 64,000 tokens |
| Images | yes, it has a vision encoder | yes |
| Licence | Apache 2.0 | Apache 2.0 |
Cloudflare's accuracy figures, on its own benchmark runs:
- BFCL case exact: Clef 98.47%, Clef-flash 98.76%
- BANKING77 macro-F1: Clef 94.20%
- CLINC150+OOS macro-F1: Clef 97.43%
The detail worth pausing on: Clef-flash, the smaller model, scores slightly higher on BFCL than the larger Clef while answering more than five times faster. On narrow, well-defined tasks, bigger is not automatically better, which is a useful corrective to the usual instinct.
Both are built on frozen Qwen backbones with rank-256 low-rank adapters — in other words, Cloudflare did not train a model from scratch. It took an open model, froze it, and trained a small set of extra weights on top to change the shape of the output.
Why is it so much faster?
209 milliseconds against a typical language model's second or more is not a tuning win. It comes from doing something structurally different.
An ordinary language model is autoregressive: it produces one token, feeds it back in, produces the next, and repeats until it stops. The answer is built a piece at a time, so a longer answer takes longer, and you pay that cost even when the answer is one word out of eleven.
Clef is non-autoregressive. It does not generate token by token at all. Cloudflare describes a "two-stage attention routing process" that arrives at a decision in one pass. Nothing is being composed, so nothing has to be composed in sequence.
Cloudflare's own comparison point is an earlier model, Jev, at 524.1 ms median with a 32,000-token context and no image support. Clef is roughly 2.5 times faster with twice the context, and Clef-flash is more than ten times faster.
If the mechanics of token-by-token generation are unfamiliar, from logits to tokens covers what a normal model is doing at each step, which is exactly what Clef skips.
What would you actually use it for?
Anything where the answer is one of a known set:
- Routing a support ticket to the right team, which is Cloudflare's own headline example.
- Deciding whether to escalate to a human, using the probability rather than the answer.
- Classifying content — spam or not, which language, which category, safe or unsafe.
- Choosing a tool in an agent loop. An agent deciding which of six tools to call is making a classification, not writing an essay. Our lesson on routing and text to SQL covers that pattern, and agentic RAG shows where it sits in a larger system.
- Triaging images, since both models have a vision encoder.
In an agent, these decisions happen constantly, and each one is latency the user feels. Shaving a second off every routing step compounds in a way that a faster final answer does not.
Clef versus just using a cheap LLM
This is the honest comparison, and it got harder for Clef in the same month it launched. Anthropic's Claude Haiku 5.5, released on 7 October 2026, is positioned for exactly these jobs — Anthropic's words are "high-volume, latency-sensitive tasks such as classification, extraction, and routing" — at a price that starts at $0.10 per million input tokens.
So why use a decision model?
| Clef | A small general LLM | |
|---|---|---|
| Output shape | one of your options, guaranteed | free text you have to parse and validate |
| Confidence | a probability you can threshold | none, beyond asking it and hoping |
| Latency | 39–209 ms | typically hundreds of ms to seconds |
| Flexibility | only classification | anything |
| Running it yourself | Apache 2.0, download the weights | depends on the model |
The honest answer is that a general model is more flexible and a decision model is more dependable. If your job genuinely is "pick one of these", the guarantees are worth more than the flexibility, and you get the probability for free. If the job is fuzzier, or you would rather maintain one model than two, the small general model wins.
How to use Clef
Two routes:
- Workers AI, Cloudflare's own API, if you want it hosted and have no wish to run a 27-billion-parameter model yourself.
- Hugging Face, where the weights are published under Apache 2.0, if you want it on your own hardware. Clef-flash at 9B is small enough to run on modest hardware; our GPU guide covers what that takes.
Common questions
What is Cloudflare Clef?
Clef is a decision model from Cloudflare, released on 1 October 2026. Rather than writing an answer in sentences, it picks one option from a list you supply and returns it with a probability. Cloudflare published two versions, Clef and the smaller Clef-flash, both under the Apache 2.0 licence.
What is a decision model?
A model whose output is constrained to a fixed set of options rather than free text, returned with a confidence score. It is built for classification, routing and tool selection, where the answer is one of a known list and a paragraph would have to be parsed back into that list anyway.
How fast is Cloudflare Clef?
Cloudflare reports a median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash, against 524.1 ms for its earlier Jev model. The speed comes from being non-autoregressive: Clef does not generate an answer token by token, so it does not pay the cost of composing text.
Is Cloudflare Clef free?
The model weights are published under the Apache 2.0 licence, so you can download and run them yourself at no licence cost. Using Clef through Cloudflare's Workers AI is a paid hosted service, priced like Cloudflare's other inference products.
What model is Clef built on?
Clef uses a frozen Qwen 3.8-27B backbone and Clef-flash a frozen Qwen 3.5-9B, with rank-256 low-rank adapters trained on top. Cloudflare did not train a model from scratch: it adapted open models to produce constrained decisions instead of free text.
Should I use Clef or a small language model like Claude Haiku?
Use Clef when the answer is genuinely one of a fixed list and you want that guaranteed, plus a probability you can set a threshold on. Use a small general model when the task is fuzzier, or when you would rather maintain one model for several jobs and accept parsing its text output.
Links
- Cloudflare: Clef decision models — 1 October 2026
- Cloudflare Workers AI
- Related: Claude models compared · agentic RAG