All posts
Decision ModelsOpen SourceSelf-HostingAI Architecture

Cloudflare Clef: Decision Models for AI Workflows – with Self-Hosting

4 min readThomas Stermole
Cloudflare Clef evaluates predefined answer options for AI workflows and can run on Workers AI or on your own infrastructure.

Cloudflare Clef is a family of AI decision models that evaluate inputs against predefined answer options. The results can feed directly into a workflow. One particularly interesting feature: Clef and Clef-flash are also available for deployment on your own infrastructure.

Cloudflare introduced the models on 1 October 2026. They can be used through Workers AI or downloaded from Hugging Face, providing a choice between a managed service and self-hosting. Official announcement

Sources checked: 5 October 2026. This introduction is based on Cloudflare's announcement, documentation and model cards. It does not include a hands-on test.

What does a decision model like Clef do?

With Clef, you provide a situation, questions and allowed answer options. The model evaluates those options and returns structured results with model probabilities. It does not need to generate free-form answers that your application must then parse. Cloudflare's launch documentation

For example, a customer reports that orders can no longer be completed in an online shop. A workflow could ask Clef which team should handle the request and how urgently it needs attention. Allowed answers might include “Technical”, “Billing” and “Sales”, alongside several urgency levels.

The application can use the assessment to route a ticket or initiate a review. The workflow still determines which actions are permitted. A model probability does not guarantee that a decision is correct.

Clef and Clef-flash at a glance

Cloudflare releases two variants. The model cards describe both as multimodal models: they can process images and video frames alongside text and JSON. Clef model card, Clef-flash model card

VariantBase modelCloudflare's positioning
ClefQwen3.8-27B, 27 billion parametersMore demanding decisions
Clef-flashQwen3.5-9B, 9 billion parametersApplications requiring particularly short response times

According to Cloudflare, the interface is compatible with Jev and SystemOne. Supported question types include yes/no questions, selection from named options and assessments using ordered answer options. Interface and question types

For more examples of the underlying approach, see Jev explained: AI Decision Models.

Self-hosting Clef: why it matters

Yes, you can self-host Clef and Clef-flash. Cloudflare provides the model weights, specialised decision head and required inference code on Hugging Face under Apache 2.0. Running the models on your own infrastructure is an explicitly supported deployment option. Clef: files and usage, Clef-flash: files and usage

For businesses, this adds another choice of operating model. Decision processing can take place in your own server environment or a private cloud you manage. This deployment route does not require a Cloudflare model endpoint.

That is particularly relevant when processing internal information or when a business wants to control infrastructure, access and model versions itself. Access to the model weights also reduces dependence on the availability of a single hosted service.

Self-hosting requires a suitable inference environment. Cloudflare's model cards document use with PyTorch, Transformers and a custom model implementation. This does not establish that installation is straightforward in every local LLM app. Hardware requirements and operating effort need to be assessed for the intended use. Documented local inference

What applications are possible?

Clef is particularly interesting where an application needs to choose between predefined options. Potential tasks include routing support requests, classifying documents or selecting the next processing step.

In a document review workflow, for example, the model could be asked whether an image is sufficiently legible and which document type it represents. Extracting the contents and processing them further would remain separate steps. These are possible use cases, not results I have tested.

Two ways to get started

Clef and Clef-flash are available on Workers AI as a managed service. For deployment on your own infrastructure, the official Hugging Face repositories are the starting point. Workers AI introduction

In my view, this choice is what makes Clef interesting: businesses can use a specialised decision model and move its processing onto their own infrastructure when needed. For teams already considering local decision models, this adds another concrete option. The article on local Jev alternatives offers further context.

Frequently asked questions about Cloudflare Clef

What is Cloudflare Clef?

Clef is a family of AI decision models from Cloudflare. The models evaluate an input against predefined questions and answer options, returning structured outputs with model probabilities.

Can you self-host Cloudflare Clef?

Yes. Cloudflare publishes Clef and Clef-flash, including model weights and inference code, on Hugging Face under Apache 2.0. This makes deployment on your own infrastructure possible; suitable hardware and a configured inference environment are required.

How does Clef differ from Clef-flash?

Clef is based on a Qwen model with 27 billion parameters, while Clef-flash uses a model with 9 billion. Cloudflare positions Clef for more demanding decisions and Clef-flash for applications that need particularly short response times.

Official sources

Next step

Sounds relevant for your company?

In a no-obligation initial call, we clarify within 30 minutes whether and where getting started is worthwhile for you — honestly and without sales pressure.

Request an initial call