Editorial illustration for Cloudflare's Clef AI Models Top 7 of 10 Decision-Making Benchmarks
Cloudflare's Clef AI Models Win 7 of 10 Benchmarks
Cloudflare's Workers AI team shipped its first two models this week, and neither one writes a single word of text. Clef and Clef-flash, released open-weight under Apache 2.0, don't generate prose or code. They answer yes/no questions, pick from named options, or score against a rubric, each time returning a typed probability instead of a sentence that needs parsing afterward. Both are live on Workers AI now, with weights posted to Hugging Face for anyone who wants to self-host.
The models plug into TypeSafe AI's Jev API, the "System One" format that debuted September 15, 2026 and has already spawned open alternatives like Kev-9B and Laya. Cloudflare built Clef to speak the same protocol, so teams running Jev today can swap in Clef by changing an endpoint and a model name, nothing more.
Under the hood, Clef is post-trained from Qwen3.8-27B and Clef-flash from Qwen3.5-9B, both keeping their backbone's vision encoder intact. A single request can bundle up to 64 questions and 4 images. The architecture behind how it turns a prefill pass into per-question probabilities is where things get specific.
Cloudflare has released Clef and Clef-flash, the first models trained by its Workers AI team. They are decision models, not chatbots. Each reads an input state and a schema of typed questions.
Why this matters For developers wiring classification logic into production systems, a model that outputs calibrated probabilities against a fixed schema is a different tool than a chatbot, and Cloudflare shipping it with open weights on Hugging Face under Apache 2.0 means teams can self-host rather than lease the capability. The jump from 79.74 to 94.20 macro-F1 on BANKING77, and 89.27 to 97.43 on CLINC150+OOS, is the kind of gain that could let intent-classification pipelines drop a fine-tuned LLM layer entirely. But three of ten benchmarks in Cloudflare's own Decision Index suite still favor Jev, and Cloudflare built both the benchmark and the model being tested against a format it's also compatible with.
That's worth flagging before anyone treats the leaderboard as neutral. Founders evaluating this for routing, moderation, or triage systems should pull the weights and run their own schema-specific tests rather than taking the 7-of-10 headline at face value. The real signal here isn't the score, it's that a major infra provider is betting decision models deserve their own architecture, separate from text generation entirely.
Common Questions Answered
What are Clef and Clef-flash models, and how do they differ from traditional chatbots?
Clef and Clef-flash are decision models released by Cloudflare's Workers AI team that don't generate text like chatbots do. Instead, they answer yes/no questions, pick from named options, or score against a rubric by returning typed probabilities rather than sentences that need parsing. This makes them specifically designed for classification logic in production systems rather than conversational AI.
What are the performance improvements of Clef models on decision-making benchmarks?
Clef models achieved significant performance gains on key benchmarks, jumping from 79.74 to 94.20 macro-F1 on BANKING77 and from 89.27 to 97.43 on CLINC150+OOS. These substantial improvements in decision-making accuracy could allow intent-classification pipelines to drop fine-tuned models and rely on the new Clef capabilities instead.
Under what license are Clef and Clef-flash released, and where can developers access them?
Clef and Clef-flash are released as open-weight models under the Apache 2.0 license, with weights posted to Hugging Face for anyone who wants to self-host. Both models are live on Workers AI, allowing developers to either use them directly through Cloudflare's service or deploy them independently in their own infrastructure.
What is the advantage of Clef models returning typed probabilities instead of text output?
By returning typed probabilities against a fixed schema, Clef models eliminate the need to parse and interpret natural language text, making them more reliable for production systems that require structured decision outputs. This approach provides calibrated probabilities that can be directly integrated into classification pipelines without additional post-processing steps.
Further Reading
- Introducing Clef: Cloudflare's first open-source decision models - Cloudflare Docs
- our open-source decision models, and new RL fine-tuning ... - Cloudflare Blog
- Cloudflare tries to outplay Jev with open-weight Clef models - The Register
- Cloudflare Adds Its First Team-Trained Decision Models to Workers AI - TokenPost
- Cloudflare Open Sources Clef to Challenge OpenAI and Amazon on AI Agent Decisions - Startup Fortune