Skip to main content
Cloudflare Clef AI model decision-making, digital interface, fast processing, cybersecurity, network optimization.

Editorial illustration for Cloudflare's Clef AI Model Makes Decisions in 39 Milliseconds

Cloudflare's Clef AI Makes Decisions in 39ms

• 4 min read

Cloudflare put a number on how fast an AI agent can make up its mind: 39 milliseconds. That's the median response time for Clef-flash, one of two new decision models the company released this week, with its larger sibling Clef clocking in at 209 milliseconds. Both are built on Qwen, handle text and images, and skip the part where a language model writes out paragraphs of reasoning. Instead, they spit out probabilities across a set of predefined answers, fast enough for other software to act on immediately.

The target here is TypeSafe AI's Jev, which has been the default choice for this kind of work. Cloudflare's models compete on the same turf: classifying a support ticket's urgency, flagging which team should handle it, deciding whether a case needs a person at all. The company frames this as removing a bottleneck rather than removing judgment, letting systems route, escalate, or defer automatically based on a score instead of waiting on a human to read and decide.

The name itself points to how Cloudflare sees the job these models do.

According to Cloudflare, "a human does not necessarily need to be in the loop for agentic decisions anymore." Agents can "programmatically gather context, make decisions, and take actions on tasks, or defer to a human when needed."

Why this matters

Cloudflare picking a fight with TypeSafe AI over decision models, down to copying Jev's API, tells us the race for AI agent infrastructure is moving past the chatbot layer. Developers building agents don't need another model that writes paragraphs; they need something that picks option A over option B in under 40 milliseconds and doesn't choke the pipeline. If Clef-flash really holds up across those 43 benchmarks, that's a meaningful speed floor for anyone running agents at scale, where latency compounds fast across chained calls.

The framing that humans "no longer need to be in the loop" is worth pushing back on. A model assigning probabilities to predefined answers is only as safe as whoever defined those answers, and Cloudflare's own pitch is automation of judgment calls that used to need a person watching. Founders evaluating this should ask what happens when the predefined options don't cover the actual situation. Worth watching: how TypeSafe responds, and whether independent benchmarks confirm Cloudflare's numbers once Clef sees real production traffic outside controlled tests.

Common Questions Answered

What is the median response time for Cloudflare's Clef-flash model?

Cloudflare's Clef-flash model has a median response time of 39 milliseconds, making it significantly faster than its larger sibling Clef, which operates at 209 milliseconds. This speed is crucial for AI agents that need to make decisions quickly without slowing down software pipelines.

How do Cloudflare's Clef decision models differ from traditional language models?

Unlike traditional language models that write out paragraphs of reasoning, Clef models skip that step and instead output probabilities across a set of predefined answers. This approach allows other software to act on the decisions immediately without processing lengthy text explanations.

What capabilities do both Clef and Clef-flash models support?

Both Clef and Clef-flash models are built on Qwen and handle both text and images as inputs. They are designed to enable AI agents to programmatically gather context, make decisions, and take actions on tasks, or defer to humans when needed.

Why does Cloudflare's focus on decision models represent a shift in AI infrastructure?

Cloudflare's Clef models indicate that the race for AI agent infrastructure is moving beyond the chatbot layer toward specialized decision-making systems. Developers building agents need models that can quickly choose between predefined options rather than generate lengthy text, which is why Clef-flash's sub-40-millisecond performance is meaningful for running agents at scale.

LIVE20:50Cloudflare's Clef AI Model Makes Decisions in 39 Milliseconds