# Cloudflare Clef
Clef is a family of two [[Decision Models (DMs)|decision models]] that [[Cloudflare]] released on 1 October 2026: **Clef** (27B) and **Clef-flash** (9B). They're the first models trained by the [[Cloudflare Workers AI]] team itself. Both are hosted on Workers AI, and the weights are on Hugging Face under the Apache 2.0 license.
Why should you care? Two weeks after [[TypeSafe AI]] launched [[Jev]], the most interesting thing about decision models was that the best one was closed, API-only and behind a waitlist. Clef changes that. It speaks the same API as Jev, so switching is a matter of changing the endpoint and the model name, and you can download the weights and run them yourself.
Like Jev, Clef never writes text. You send it some state plus up to 64 typed questions, and it returns a probability for every allowed answer. Your agent gets a decision it can act on right away (route the ticket, block the request, escalate to a human), with nothing to parse and no reasoning tokens to wait for.
## What it adds on top of Jev
- **Vision.** Clef has a vision encoder. On Workers AI you can attach up to 4 images (PNG, JPEG or WebP, embedded in the request; remote URLs aren't accepted). The Hugging Face cards say the model also reads video frames when you run it yourself. Jev only handles text today
- **A bigger context window.** 64K tokens, against 32K for Jev
- **Open weights.** Jev's weights, architecture and training are all private. Clef's weights are public, and Cloudflare describes the architecture and the training recipe (see below)
- **Edge hosting.** Requests run on GPUs across Cloudflare's network, close to your users. The idea is to put Clef directly in your agent's request path, then hand off to an LLM on Workers AI when something needs to happen
The question types are the [[System One Primitives]]: `noul` (yes/no, returns the probability of yes), `choice` (one option from a set you define, with a probability per option and a confidence) and `score` (a probability-weighted score on an ordered rubric). You can call it through the Workers AI binding (`env.AI.run("@cf/cloudflare/clef", ...)`), through the REST API at `/ai/run`, or through [[Cloudflare AI Gateway]].
## Numbers
Everything here comes from Cloudflare, so treat it as vendor benchmarks.
- **Latency** (median / p95, across 43 benchmark runs): Clef 209.3 / 238.6 ms, Clef-flash 38.8 / 122.4 ms, Jev 524.1 / 536.0 ms. So Clef is about 2.5x faster than Jev at the median, and Clef-flash about 13x. Note that Jev was measured as a hosted API (with a network round trip), and that some open models are faster still ([[Laya]] has a 5.8 ms median, but it scores far lower)
- **Benchmarks:** a Clef model comes out on top on 7 of the 10 decision benchmarks Cloudflare picked. Examples: BANKING77 macro-F1 94.20 (Clef) vs 79.74 (Jev), CLINC150+OOS 97.43 vs 89.27, Home appliances 97.73 (Clef-flash) vs 52.27. Jev still wins on When2Call (80.97 vs 72.37) and BRIGHT (47.52 vs 45.91), and a DiffusionGemma-based Jev reproduction leads on PhishNChips (85.35 vs 79.60)
- **TypeSafe's own workflow evals:** Clef beats Jev on invoice processing (64.7 vs 61.8), customer service (76.3 vs 76.0) and security incidents (62.9 vs 61.7), and loses on agent trace observability (68.5 vs 71.6). Most of those gaps are small
- **Decision Index 0.2.1 leaderboard:** Clef ranks first at 61.21, ahead of Jev at 57.91, with Clef-flash at 57.07. But the Clef rows are marked "self-reported", cover 36 of the 38 benchmarks, and come with no calibration figure (Jev has an ECE of 0.074 on the same board)
- **A real workflow:** Cloudflare's Threat Intelligence team uses Clef with [[Cloudflare Browser Rendering|Browser Run]] to classify domains (e.g., 95% fashion, 85% ecommerce, under 1% phishing). Fetching, rendering and classifying took 2.2 s, against 4.7 s for gpt-oss-120b in the same workflow, which also returned only two categories
- **Price:** $0.24 per million input tokens for Clef, $0.09 for Clef-flash. That's MORE expensive than Jev's $0.042, about 6x for the large model, as an HN commenter pointed out
## How it was trained
TypeSafe never published any of this for Jev. Cloudflare did.
- **Backbone.** A frozen [[Qwen]] model: Qwen3.8-27B for Clef, Qwen3.5-9B for Clef-flash, vision encoder included
- **Joint schema head.** A small transformer head reads the backbone's final hidden states, routes evidence from the state to each question, lets the questions attend to each other, and scores every option of every question jointly. Inference is one prefill pass followed by parallel scoring; nothing is generated token by token
- **Adapters and losses.** The routing head was trained jointly with rank-256 LoRA adapters, using label-smoothed cross-entropy for valid outputs plus a Brier loss to improve calibration ([[Proper Scoring Rules]], [[AI Model Calibration]])
- **Data.** Cloudflare's own synthetic datasets, with field orders, prompts and schema structures shuffled
- **RLCD.** A second optimization stage that Cloudflare also calls [[Reinforcement Learning for Calibrated Decisions (RLCD)|Reinforcement Learning for Calibrated Decisions]] (the name TypeSafe introduced). It gives partial credit to adjacent ordinal choices, rewards fully correct records, and adds a reference penalty against distribution shift
It didn't come out of nowhere. On 18 September, three days after Jev launched, Michelle Chen of Cloudflare posted a demo that simulated Jev with an open LLM ([[DiffusionGemma]]) and its token probabilities, running on Workers AI. That work built on Matt Mastracci's [[vLLM]] pull request. Clef keeps the idea but swaps in a Qwen backbone and a trained head.
## The RL fine-tuning service
Cloudflare launched a fine-tuning service at the same time. It starts as hands-on work with their forward-deployed engineers (design partners sign up through a form) and should become a self-serve platform later. The pipeline reuses existing pieces: AI Gateway to capture your traffic as a dataset, Workers AI for rollouts against the base model, [[Cloudflare Containers]] as RL sandboxes, a new "Trainer" to update the weights, and Workers AI's bring-your-own-model support to redeploy the result. Internally, they plan to tune Clef for Trust & Safety reviews, support triage and good-bot/bad-bot decisions.
## What the community said
The Hacker News thread took off (623 points, 215 comments on launch day). The reactions worth keeping:
- **"No mention of calibration."** For a decision model, that's the main property, and the leaderboard has no ECE for Clef yet
- **Open weights, not open source.** The weights are Apache 2.0, but the training data and pipeline aren't published, so nobody can reproduce them
- **Price.** The large model is ~6x Jev's price; Clef-flash is the competitive option
- **Speed of cloning.** Several people asked how so many Jev alternatives showed up within weeks, and some felt bad for TypeSafe. The leaderboard already lists more than 70 entries
## My take
I think Clef matters more for what it proves than for its benchmark lead. A public company shipped an open-weight decision model that matches or beats Jev on most published numbers, two weeks after Jev launched. So the interface (state in, typed probabilities out) is now a commodity, and the moat has to be somewhere else. My bet is on calibration and training data, plus whoever already sits in the request path (which is exactly where Cloudflare is).
Would I use it? For triage and routing on Cloudflare, Clef-flash looks like the obvious pick: 38.8 ms at the median, vision support, and no waitlist. But I wouldn't take "it beats Jev" at face value. The numbers are self-reported, Clef is pricier per token, and nobody has published calibration figures for it yet. As with Jev, test it on your own data before you trust a threshold.
## References
- [Introducing Clef: Cloudflare's first open-source decision models, now on Workers AI (Cloudflare changelog, 2026-10-01)](https://developers.cloudflare.com/changelog/post/2026-10-01-clef-workers-ai/)
- [Introducing Clef: our open-source decision models, and new RL fine-tuning platform (Cloudflare blog)](https://blog.cloudflare.com/clef-decision-models/)
- [Clef model page (Workers AI docs)](https://developers.cloudflare.com/workers-ai/models/clef/)
- [Clef-flash model page (Workers AI docs)](https://developers.cloudflare.com/workers-ai/models/clef-flash/)
- [Cloudflare/clef (Hugging Face model card)](https://huggingface.co/Cloudflare/clef)
- [Cloudflare/clef-flash (Hugging Face model card)](https://huggingface.co/Cloudflare/clef-flash)
- [Decision Model Leaderboard (Cloudflare, data from Decision Index 0.2.1)](https://clef-evals.workers-ai-mle.workers.dev/)
- [Jev Decision Index (Hugging Face Space)](https://huggingface.co/spaces/multimodalart/jev-decision-index)
- [Introducing System One Models & Jev (TypeSafe AI)](https://typesafe.ai/blog/introducing-system-one-models-and-jev)
- [Michelle Chen's DiffusionGemma decision-model demo (X, 2026-09-18)](https://x.com/michellechen/status/2101091012559151480)
- [vLLM pull request 57250 (Matt Mastracci)](https://github.com/vllm-project/vllm/pull/57250)
- [Clef RL fine-tuning design partner form](https://www.cloudflare.com/resource/clef-rl-interest)
- [Workers AI pricing](https://developers.cloudflare.com/workers-ai/platform/pricing/)
- [Hacker News discussion](https://news.ycombinator.com/item?id=49923692)
## Related
- [[Jev]]
- [[Decision Models (DMs)]]
- [[System One Models]]
- [[System One Primitives]]
- [[TypeSafe AI]]
- [[Kev]]
- [[Laya]]
- [[DiffusionGemma]]
- [[Cloudflare Workers AI]]
- [[Cloudflare]]
- [[Cloudflare AI Gateway]]
- [[Cloudflare Agents SDK]]
- [[Qwen]]
- [[Reinforcement Learning for Calibrated Decisions (RLCD)]]
- [[AI Model Calibration]]
- [[AI Open Weight Models]]