Cloudflare Clef: Decision Models for AI Workflows – with Self-Hosting
Cloudflare Clef is a family of AI decision models that evaluate inputs against predefined answer options. The results can feed directly into a workflow. One particularly interesting feature: Clef and Clef-flash are also available for deployment on your own infrastructure.
Cloudflare introduced the models on 1 October 2026. They can be used through Workers AI or downloaded from Hugging Face, providing a choice between a managed service and self-hosting. Official announcement
Sources checked: 5 October 2026. This introduction is based on Cloudflare's announcement, documentation and model cards. It does not include a hands-on test.
What does a decision model like Clef do?
With Clef, you provide a situation, questions and allowed answer options. The model evaluates those options and returns structured results with model probabilities. It does not need to generate free-form answers that your application must then parse. Cloudflare's launch documentation
For example, a customer reports that orders can no longer be completed in an online shop. A workflow could ask Clef which team should handle the request and how urgently it needs attention. Allowed answers might include “Technical”, “Billing” and “Sales”, alongside several urgency levels.
The application can use the assessment to route a ticket or initiate a review. The workflow still determines which actions are permitted. A model probability does not guarantee that a decision is correct.
Clef and Clef-flash at a glance
Cloudflare releases two variants. The model cards describe both as multimodal models: they can process images and video frames alongside text and JSON. Clef model card, Clef-flash model card
| Variant | Base model | Cloudflare's positioning |
|---|---|---|
| Clef | Qwen3.8-27B, 27 billion parameters | More demanding decisions |
| Clef-flash | Qwen3.5-9B, 9 billion parameters | Applications requiring particularly short response times |
According to Cloudflare, the interface is compatible with Jev and SystemOne. Supported question types include yes/no questions, selection from named options and assessments using ordered answer options. Interface and question types
For more examples of the underlying approach, see Jev explained: AI Decision Models.
Self-hosting Clef: why it matters
Yes, you can self-host Clef and Clef-flash. Cloudflare provides the model weights, specialised decision head and required inference code on Hugging Face under Apache 2.0. Running the models on your own infrastructure is an explicitly supported deployment option. Clef: files and usage, Clef-flash: files and usage
For businesses, this adds another choice of operating model. Decision processing can take place in your own server environment or a private cloud you manage. This deployment route does not require a Cloudflare model endpoint.
That is particularly relevant when processing internal information or when a business wants to control infrastructure, access and model versions itself. Access to the model weights also reduces dependence on the availability of a single hosted service.
Self-hosting requires a suitable inference environment. Cloudflare's model cards document use with PyTorch, Transformers and a custom model implementation. This does not establish that installation is straightforward in every local LLM app. Hardware requirements and operating effort need to be assessed for the intended use. Documented local inference
What applications are possible?
Clef is particularly interesting where an application needs to choose between predefined options. Potential tasks include routing support requests, classifying documents or selecting the next processing step.
In a document review workflow, for example, the model could be asked whether an image is sufficiently legible and which document type it represents. Extracting the contents and processing them further would remain separate steps. These are possible use cases, not results I have tested.
Two ways to get started
Clef and Clef-flash are available on Workers AI as a managed service. For deployment on your own infrastructure, the official Hugging Face repositories are the starting point. Workers AI introduction
In my view, this choice is what makes Clef interesting: businesses can use a specialised decision model and move its processing onto their own infrastructure when needed. For teams already considering local decision models, this adds another concrete option. The article on local Jev alternatives offers further context.
Frequently asked questions about Cloudflare Clef
What is Cloudflare Clef?
Clef is a family of AI decision models from Cloudflare. The models evaluate an input against predefined questions and answer options, returning structured outputs with model probabilities.
Can you self-host Cloudflare Clef?
Yes. Cloudflare publishes Clef and Clef-flash, including model weights and inference code, on Hugging Face under Apache 2.0. This makes deployment on your own infrastructure possible; suitable hardware and a configured inference environment are required.
How does Clef differ from Clef-flash?
Clef is based on a Qwen model with 27 billion parameters, while Clef-flash uses a model with 9 billion. Cloudflare positions Clef for more demanding decisions and Clef-flash for applications that need particularly short response times.