A deep dive into Clef, Cloudflare's decision model
By Flavio Copes
Clef is Cloudflare's open-weight decision model on Workers AI. How it works, how to call it, and how it compares with Jev and OpenAI's Decisions API.
Clef is a decision model from Cloudflare. You send it some data and a few typed questions, and it sends back probabilities: how likely the answer is yes, which option from your list fits best, where the data sits on a scale you described. It doesn’t write text.
Cloudflare released it on October 1, 2026, in two sizes. Clef is a 27 billion parameter model for the most accurate answers, and Clef-flash is a 9 billion parameter model for answers in a few tens of milliseconds. Both run on Workers AI, and the weights are on Hugging Face under the Apache 2.0 license, so you can download them and run them on your own hardware.
Clef uses the same request and response format as Jev, the System One model TypeSafe launched on September 15, so most code written for Jev works with Clef. My deep dive into Jev covers this kind of model in more depth.
Clef is the third decision model in two weeks. TypeSafe launched Jev on September 15, OpenAI announced its Decisions API in limited preview on September 29, and Cloudflare shipped Clef two days later.
I ran every example in this post on my Cloudflare account with both models, and the outputs you’ll see are the real ones.
What a decision model does
A regular if works when your code can compute the condition, like if (order.total > 100). It falls apart when the condition needs understanding. Is this support message urgent? Which team should handle it? Does this screenshot show a password?
The usual answer is to ask an LLM and parse what it writes back. A decision model skips the writing. You define the possible answers up front, the model gives each one a probability, and your code reads the numbers and branches.
There are three question types, with the names TypeSafe picked for Jev:
| Type | The question | What comes back |
|---|---|---|
noul | Is this true? | noul, a probability from 0 to 1 |
choice | Which of these options? | choice, probabilities, confidence |
score | Where on this scale? | score, legend, probabilities, confidence |
A Choice picks one option from a set of 2 to 255 options you define. A Score places the data on an ordered scale of 2 to 10 levels you describe, lowest first, and returns a probability-weighted level that can land between two of them. You can ask up to 64 questions in one request, and they all look at the same state, the data you send.
How Clef works
Cloudflare built Clef on top of Qwen, Alibaba’s open model family. Clef starts from Qwen3.8-27B and Clef-flash from Qwen3.5-9B, and both keep Qwen’s vision encoder, which is why Clef can look at images.
An LLM answers by writing, one token at a time. Clef doesn’t. The Qwen model reads your state and questions once, the same step an LLM performs before it starts writing, which Cloudflare calls a prefill-only pass. Then a small extra network on top, the joint schema head, looks at what Qwen understood and scores every allowed option of every question in one go. A softmax turns those scores into probabilities.
Since nothing is generated, an answer can’t contain an option you didn’t define, and there’s no JSON to repair. It’s also why it’s fast: one pass through the model instead of a loop.
For training, Cloudflare froze the Qwen weights and trained the head together with low-rank adapters (LoRA) on synthetic datasets that shuffle field orders, prompts and schemas. The loss rewards both correct answers and well-calibrated probabilities. On top of that there’s a reinforcement learning stage Cloudflare calls Reinforcement Learning for Calibrated Decisions (RLCD), the same name TypeSafe uses for the method behind Jev, which gives partial credit when a Score answer lands on the level next to the right one.
The name comes from music. A clef is the symbol at the start of a staff that tells you which note each line stands for, and Cloudflare sees a decision model playing a similar role for the actions that follow it. The “CF” in the middle stands for Cloudflare.
Clef or Clef-flash?
| Clef | Clef-flash | |
|---|---|---|
| Model ID | @cf/cloudflare/clef | @cf/cloudflare/clef-flash |
| Size | 27B, from Qwen3.8-27B | 9B, from Qwen3.5-9B |
| Context window | 65,536 tokens | 65,536 tokens |
| Images | Up to 4 per request | Up to 4 per request |
| Price per million input tokens | $0.24 | $0.09 |
| Median latency in Cloudflare’s tests | 209 ms | 39 ms |
Cloudflare positions Clef as the precision model and Clef-flash as the one for the hot path, where a request waits for the answer. The prices come from the Workers AI pricing page as of October 1, 2026. Output isn’t billed, and every response I got reported "output_tokens": 0.
My advice is to start with Clef-flash and move a question to Clef only when your labeled examples show Flash getting it wrong. The first example shows why.
Your first call with curl
You need two values from the Cloudflare dashboard. Open Workers AI, select Use REST API, and create a token with the Create a Workers AI API Token template. Copy your account ID from the same page. If you make the token by hand instead, it needs both Workers AI - Read and Workers AI - Edit.
Save them in the environment, with the names the Cloudflare docs use:
export CLOUDFLARE_ACCOUNT_ID=your_account_id
export CLOUDFLARE_AUTH_TOKEN=your_api_token
Now let’s ask Clef-flash three questions about a support ticket. It’s the same ticket I used in the Jev deep dive, so you can compare the answers:
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef-flash \
-X POST \
-H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
-d '{
"model": "clef-flash",
"state": {
"ticket": "The export button crashes the settings page in Safari. Works in Chrome, but some of our customers only use Safari."
},
"questions": {
"category": {
"type": "choice",
"instructions": "What kind of ticket is `ticket`?",
"criteria": {
"bug_report": "Something is broken or behaving wrong",
"feature_request": "Asks for something that does not exist yet",
"billing": "Charges, invoices, refunds",
"other": null
}
},
"severity": {
"type": "score",
"instructions": "How severe is the issue in `ticket`?",
"criteria": [
"Cosmetic; no impact on functionality",
"Broken or degraded feature, but a workaround exists",
"Blocking issue; no workaround exists"
]
},
"has_repro_steps": {
"type": "noul",
"instructions": "Does `ticket` say how to reproduce the problem?"
}
}
}'
The model name appears twice, in the URL and in the model field. The field is required, and it only accepts clef or clef-flash.
This is what came back:
{
"result": {
"model": "clef-flash",
"answers": {
"category": {
"type": "choice",
"choice": "bug_report",
"probabilities": {
"bug_report": 0.9664,
"feature_request": 0.0135,
"billing": 0.0058,
"other": 0.0143
},
"confidence": 0.9126
},
"severity": {
"type": "score",
"score": 1.1069,
"legend": {
"0": "Cosmetic; no impact on functionality",
"1": "Broken or degraded feature, but a workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": { "0": 0.0636, "1": 0.766, "2": 0.1704 },
"confidence": 0.4298
},
"has_repro_steps": { "type": "noul", "noul": 0.5355 }
},
"usage": { "input_tokens": 395, "output_tokens": 0 }
},
"success": true,
"errors": [],
"messages": []
}
Notice the outer object. The REST API wraps the answer in Cloudflare’s usual envelope, with success, errors and messages, so the Jev-shaped part lives under result. Also, model says clef-flash, with no version number.
Then I sent the same request to Clef. Both models are sure it’s a bug report, and they disagree on the other two questions:
| Question | Clef | Clef-flash |
|---|---|---|
category | bug_report, confidence 0.91 | bug_report, confidence 0.91 |
severity | 1.38, confidence 0.22 | 1.11, confidence 0.43 |
has_repro_steps | 0.04 | 0.54 |
Clef reads the severity as split between “a workaround exists” and “blocking”, which is fair: Chrome is a workaround, except for the customers who only use Safari. Flash leans toward the workaround.
The repro question is where they really differ. Clef says no, while Flash sits at 0.54, which means it can’t tell. I can’t tell either. “The export button crashes the settings page in Safari” is close to a reproduction step, but it’s not a list of steps. A vague question gets a vague answer, and the fix is to say what counts: “Does ticket name the browser and the action that trigger the problem?”
When an answer sits near 0.5, or the confidence is low, read your question again before you blame the model.
Use it from a Worker
Inside a Worker, the Workers AI binding takes the place of the token. Add it to wrangler.jsonc:
{
"name": "ticket-router",
"main": "src/index.ts",
"compatibility_date": "2026-09-01",
"ai": { "binding": "AI" }
}
Then you call env.AI.run() with the same body. Here’s a small Worker that receives a support message, asks Clef-flash two questions, and decides where the message goes:
type TicketAnswers = {
urgent: { noul: number }
team: { choice: 'billing' | 'technical' | 'sales'; confidence: number }
}
export default {
async fetch(request, env): Promise<Response> {
const { message } = await request.json<{ message: string }>()
const { answers } = (await env.AI.run('@cf/cloudflare/clef-flash', {
model: 'clef-flash',
state: { message },
questions: {
urgent: {
type: 'noul',
instructions: 'Does `message` describe a problem that blocks customers right now?',
},
team: {
type: 'choice',
instructions: 'Which team should handle `message`?',
criteria: {
billing: 'Payments, invoices, and refunds',
technical: 'Outages, errors, and configuration',
sales: 'Plans and upgrades',
},
},
},
})) as { answers: TicketAnswers }
if (answers.team.confidence < 0.5) {
return Response.json({ route: 'human' })
}
return Response.json({
route: answers.team.choice,
page_on_call: answers.urgent.noul > 0.8,
})
},
} satisfies ExportedHandler<Env>
The binding returns the Jev-shaped object directly, with no envelope. The TicketAnswers type is there because, as of October 1, 2026, the types generated by wrangler types don’t know the Clef models yet, so answers comes back as unknown. With the type you also get autocomplete on the three teams.
I ran it with npx wrangler dev, which runs the Worker on your machine but sends the AI calls to Cloudflare, and posted three messages to it:
"Checkout has been failing for every customer for the last hour."
{"route":"technical","page_on_call":true}
"Can I switch my team from the monthly plan to the yearly one?"
{"route":"sales","page_on_call":false}
"I was charged twice this month, can you refund one?"
{"route":"billing","page_on_call":false}
The model makes the judgment, and the thresholds that decide what happens next live in your code, where you can read them and change them. If you’re new to Workers AI, my Workers AI guide covers bindings and the free allocation, and the free Cloudflare course has a whole module on it.
Switching from Jev to Clef
Cloudflare says you can switch an existing Jev integration by changing the endpoint and the model. You also change the key, and over REST the response gets a wrapper. The full list:
- The URL becomes
https://api.cloudflare.com/client/v4/accounts/<account ID>/ai/run/@cf/cloudflare/clefinstead ofhttps://api.typesafe.ai/v1/systemone. - The key becomes a Cloudflare API token with Workers AI permissions.
- The
modelfield must becleforclef-flash. If you leavejev-latestin, the request fails with error code 5006, which says the field doesn’t match that pattern. - Over REST, the answers sit under
result.
The state, the questions and the answer shapes stay the same.
If you use TypeSafe’s JavaScript SDK, you can keep your code, the noul, choice and score helpers, and the typed answers. The client accepts a custom fetch function, so you can send its requests to Workers AI and unwrap the envelope there:
import { choice, noul, score, TypeSafeClient } from '@typesafe-ai/sdk'
const account = process.env.CLOUDFLARE_ACCOUNT_ID
async function workersAi(url, init) {
const { model } = JSON.parse(init.body)
const res = await fetch(`https://api.cloudflare.com/client/v4/accounts/${account}/ai/run/@cf/cloudflare/${model}`, init)
const { success, result, errors } = await res.json()
if (!success) return Response.json({ errors }, { status: res.status })
return Response.json(result)
}
const client = new TypeSafeClient({
apiKey: process.env.CLOUDFLARE_AUTH_TOKEN,
defaultModel: 'clef-flash',
fetch: workersAi,
})
const { answers } = await client.systemOne({
state: { ticket: 'I was charged twice for order A-104. Please refund the duplicate.' },
questions: {
department: choice('Which team should handle `ticket`?', {
billing: 'Charges, invoices, refunds',
technical: 'Bugs, outages, integration problems',
other: null,
}),
refund_requested: noul('Does `ticket` ask for money back?'),
frustration: score('How frustrated is the author of `ticket`?', [
'Calm, just stating facts',
'Frustrated but civil',
'Very angry or threatening to leave',
]),
},
})
console.log(answers.department.choice)
It prints billing, with both models. The SDK still thinks it’s talking to TypeSafe, so anything other than systemOne(), like client.models.list(), won’t work through this adapter. My Jev in Node.js guide covers the rest of the SDK.
Clef can look at images
This is the feature Jev doesn’t have. Next to state, Clef accepts an images array with up to four PNG, JPEG or WebP images, sent as base64. Remote URLs aren’t accepted.
I tested it on a check I do by hand before I publish a post: what a screenshot shows, and whether it reveals something it shouldn’t. I picked two screenshots from this site. One is the dark terminal from my Cloudflare Quick Tunnels guide, which shows a tunnel URL and local IP addresses. The other is a light screenshot of Fish Audio’s API keys page, with no key on it.
import { readFileSync } from 'node:fs'
const image = readFileSync('terminal.jpg').toString('base64')
const res = await fetch(`https://api.cloudflare.com/client/v4/accounts/${process.env.CLOUDFLARE_ACCOUNT_ID}/ai/run/@cf/cloudflare/clef`, {
method: 'POST',
headers: { Authorization: `Bearer ${process.env.CLOUDFLARE_AUTH_TOKEN}` },
body: JSON.stringify({
model: 'clef',
state: 'A screenshot I want to publish in a blog post.',
images: [`data:image/jpeg;base64,${image}`],
questions: {
kind: {
type: 'choice',
instructions: 'What does the screenshot show?',
criteria: {
terminal: 'A terminal window with command output',
code_editor: 'A code editor with source code',
web_dashboard: 'A logged-in web app or dashboard',
web_page: 'A public website page',
other: null,
},
},
dark_mode: {
type: 'noul',
instructions: 'Is the main window in the screenshot using a dark theme?',
},
shows_secret: {
type: 'noul',
instructions: 'Does the screenshot show an API key, password or access token in readable form?',
},
shows_network_details: {
type: 'noul',
instructions: 'Does the screenshot show IP addresses or public URLs?',
},
},
}),
})
const { result } = await res.json()
console.log(result.answers)
Here’s what both models answered:
| Question | Terminal, Clef | Terminal, Flash | Dashboard, Clef | Dashboard, Flash |
|---|---|---|---|---|
kind | terminal, 0.94 | terminal, 0.89 | web_dashboard, 0.87 | web_dashboard, 0.83 |
dark_mode | 0.83 | 0.81 | 0.01 | 0.02 |
shows_secret | 0.07 | 0.08 | 0.02 | 0.02 |
shows_network_details | 0.92 | 0.79 | 0.67 | 0.21 |
They got everything right except one judgment call. The dashboard screenshot shows fish.audio/app/api-keys/ in the browser’s address bar. Clef noticed it at 0.67, Flash missed it at 0.21.
Image size was the first problem. My first attempt sent the original PNGs, 1.3 MB and 575 KB, well under the documented 4 MiB limit, and every request failed with exceeded this model context window limit (65536). Workers AI had estimated the terminal shot at 425,610 tokens, which is the length of its base64 string divided by four. It counts the image as text before the model sees it. Resizing both to 1024 pixels wide as JPEGs at 80% quality (140 KB and 48 KB) fixed it, and the billed input was 1,048 and 1,272 tokens. As of October 1, 2026, keep each image under about 190 KB.
Speed was the second. Text requests answered in well under a second, but every image request I sent on October 1 took between 13 and 30 seconds, with both models. That was launch day, and Cloudflare’s own threat intelligence example classified a rendered website in 2.2 seconds, so I expect this to improve. Until it does, I wouldn’t put image questions in a request a user is waiting for.
The model card on Hugging Face also mentions video frames. The Workers AI API only documents images.
What the benchmarks say
Cloudflare’s announcement has a table of 10 benchmarks where Clef beats Jev on most rows. The Hugging Face model card has the full run: 43 benchmarks plus latency, from the Decision Index, a community project on Hugging Face that benchmarks the open reproductions of Jev.
On the overall index, Cloudflare’s leaderboard puts Clef first at 61.2, Jev second at 57.9, and Clef-flash at 57.1. Two community models built on Google’s open Gemma 4 land between Jev and Clef-flash.
The full table is more mixed than the blog post’s ten rows. Clef and Clef-flash win most of the classification, routing and tool-selection benchmarks, like BANKING77, where Clef scores 94.2 and Jev 79.7. Jev still wins the reasoning-heavy ones: GPQA Diamond (78.3 against 48.0 for Clef), MMLU-Pro (82.7 against 65.9) and BBH (92.9 against 73.7). On TypeSafe’s own four workflow evals, Clef wins three, by 0.3 to 2.9 points, and Jev wins the fourth.
The leaderboard’s own methodology notes say Clef’s results are self-reported: Cloudflare ran them, and the people who maintain the index haven’t reproduced them. They also say the latency numbers aren’t comparable. Clef ran on Cloudflare’s own serving stack, while Jev’s 524 ms median is a round trip from Cloudflare’s lab to TypeSafe’s hosted API over the internet. TypeSafe quotes around 100 milliseconds for most calls, measured from the US West Coast. TypeSafe also chose not to publish public benchmark results for Jev, so all these numbers come from the community and from Cloudflare.
I measured latency myself on October 1, 2026, from Italy, through the REST API, with the 395-token ticket from the curl example. I sent ten calls in a row per model and repeated the run three times:
- Clef-flash had a median between 191 and 205 ms, and its slowest call took 676 ms.
- Clef had a median between 524 and 726 ms, with single calls up to 3 seconds.
That includes the trip from my Mac to Cloudflare and the REST API layer, so a Worker using the binding should do better. I didn’t run Jev side by side, so these are Clef numbers only.
What Clef costs
Workers AI bills in Neurons, its own unit, at $0.011 per 1,000 Neurons, and the pricing page converts that into tokens for each model. As of October 1, 2026:
| Model | Input per million tokens | Output |
|---|---|---|
| Clef | $0.24 | Not billed |
| Clef-flash | $0.09 | Not billed |
| Jev, for comparison | $0.042 | Free |
Every account gets 10,000 Neurons a day for free, on both the Free and the Paid Workers plans. That’s about 458,000 Clef tokens or 1.2 million Clef-flash tokens a day. Past that, you need the Workers Paid plan.
So Clef isn’t cheaper than Jev. Per input token, Clef costs about 5.7 times as much and Clef-flash about 2.1 times. For 100,000 tickets like the one above, at about 400 tokens each, that’s $1.68 with Jev, $3.60 with Clef-flash and $9.60 with Clef. All three are a small fraction of what a general LLM costs for the same job, which is the comparison both Cloudflare and TypeSafe care about. My Jev pricing guide has the LLM side of that math.
Downloading the weights costs nothing, running them is another story. The model card’s code loads the model with Hugging Face transformers and PyTorch, and Cloudflare tested it on a single NVIDIA H200, a datacenter GPU. That’s a research setup, not a one-command install.
Clef, Jev and OpenAI’s Decisions API
| Jev | Clef | OpenAI Decisions API | |
|---|---|---|---|
| Made by | TypeSafe | Cloudflare | OpenAI |
| Announced | September 15, 2026 | October 1, 2026 | September 29, 2026 |
| Status | Available, new signups paused since September 22 | Available on Workers AI | Limited preview |
| Model | Jev, closed | Clef and Clef-flash, open weights | GPT-6 Luna, closed |
| Input | Text | Text, plus up to 4 images | Text or images |
| API | System One | System One compatible | Not published yet |
| Price per million input tokens | $0.042 | $0.24, Flash $0.09 | Not published yet |
OpenAI’s version is the one we know least about. OpenAI says it’s powered by GPT-6 Luna, takes text or images as context, and lets you define questions with a finite set of answers to classify content, route requests or choose an agent’s next action. On September 29 it was open to selected API customers, with a broad release planned “in the coming days”. There’s no public API reference yet, so we don’t know the request format, whether it returns probabilities, or the price.
Jev has the lowest price and the most tooling around it: SDKs for JavaScript and Python, Vercel’s AI Gateway, the AI SDK’s experimental_evaluate, and plenty of community projects. Clef has open weights and image input, and it lives next to the rest of your app if that app already runs on Cloudflare.
Because Clef copied Jev’s API, code written for one now works with the other after a few changes. That makes it cheap to run both on your own data and keep the one that answers better.
How I would use Clef
I haven’t put Clef in production. Here’s where I would.
This site runs on Cloudflare Pages, and its forms already go through Pages Functions, including the sponsor form. In my Jev deep dive I planned to label sponsor inquiries with a Noul for “is this a real sponsorship request” and a Choice for the product category. With Clef I could do the same with a Workers AI binding inside the function I already have, with no new vendor and no new secret to store. I’d start by adding the labels to the email I already receive and keep deciding myself.
The second place is the screenshot check from the image example. Before a post goes live I look at every new screenshot, and I ask the same questions every time: is it light mode, does it show a key, a URL or an IP address? A script could send each new image in the post’s folder to Clef and print the ones that need a second look. Checking the image width stays in code, because that’s a number, not a judgment. And since it runs before publishing, the slow image requests don’t matter.
Where I wouldn’t use it:
- Big batch jobs where the price per token matters most, since Jev costs less.
- Questions that need reasoning more than classification, where Jev scored higher in the full benchmark table.
- Image questions in a request someone waits for, until the latency I measured on launch day comes down.
- Anything a plain
ifalready decides correctly.
The fine-tuning service
Cloudflare also announced a reinforcement learning service to fine-tune Clef on your own data. For now it’s hands-on: Cloudflare’s forward-deployed engineers work with you, and you sign up as a design partner. A self-serve version comes later.
The plan connects pieces Cloudflare already has. AI Gateway records your real requests as a dataset, Workers AI generates answers with the base model, Containers score and replay them, a new Trainer updates the weights, and Workers AI serves the result as your own model. Cloudflare says it’s working on the same thing internally for Trust & Safety reports, support triage and telling good bots from bad ones.
If your AI traffic already goes through AI Gateway, that’s where the training data would come from.
Where to start
Create a Workers AI token, copy the curl example, and replace the ticket with one real input from your own app. Then write down 20 inputs where you know the right answer, run them through Clef-flash and Clef, and look at the ones where the two models disagree or the confidence is low. Rewrite those questions first.
The Clef model page has the full schema, and Cloudflare’s launch post has the training details.
Want me to talk about your product? You can sponsor this site.
Related posts about cloudflare: