News · 2026-10-02
Cloudflare launches Clef models that score decisions without writing an answer
Cloudflare launched Clef and Clef-Flash on October 1 to score developer-defined decisions directly instead of generating conversational answers. The company released hosted models and open weights, and reported a domain-classification workflow taking 2.2 seconds with Clef versus 4.7 seconds with its chosen general-model comparison.
Key facts
- Cloudflare announced the pair on October 1, 2026.
- Its domain workflow took 2.2 seconds versus 4.7 seconds in a vendor-run demonstration.
- Hosted rates are $0.24 per million input tokens for Clef and $0.09 for Clef-Flash.
- Primary sources are Cloudflare’s launch post and official model documentation.
Many applications ask an AI a short operational question: which queue receives this message, which tool should run next, or does this domain look like phishing? A general chatbot can answer, but software then has to parse the response and handle outputs outside the expected format. Clef turns the allowed answer space into an explicit input.
Cloudflare’s interface accepts a state and a schema. The state can contain the material to evaluate; the schema defines questions and valid answers. The model returns yes-or-no probabilities, choices among named alternatives, or ordered scores. The hosted interface supports up to 64 named questions and four embedded images. Its documentation does not accept remote image URLs in that API.
The mechanism matters more than the new category label. Cloudflare says the model reads the input in a prefill-only pass, then a joint scoring head evaluates valid answers in parallel. Information can flow between question fields before scoring. Think of an examiner filling a multiple-choice answer sheet directly after reading an assignment, rather than dictating a paragraph that an assistant must transcribe and interpret.
The Clef model card identifies a Qwen3.8-27B backbone, while Clef-Flash uses Qwen3.5-9B. Both include vision components and are released under Apache 2.0. The published artifacts include backbone files, a separate joint head, configuration, and inference glue. The dossier does not verify their weight-download totals, so no disk-size estimate is supplied. Cloudflare reports testing on a single H200, a 141 GB accelerator; that is a tested configuration, not a stated minimum video-memory requirement.
Cloudflare froze the larger backbone and trained a routing head alongside low-rank adapters. It describes label-smoothed classification loss and a probability-scoring loss, plus a secondary objective named “Reinforcement Learning for Calibrated Decisions.” Those are disclosed training choices. The launch does not publish a reliability plot showing that an answer assigned 80% confidence is correct about 80% of the time.
The practical distinction is explained in our lesson on calibration. A model can choose the right class frequently while expressing unreliable confidence. That matters when application code uses a score to approve an action or decide whether a human should review it. The developer may also need to include an explicit unknown option or implement an escalation policy; a largest probability among supplied choices does not establish that any choice fits.
Cloudflare’s clearest demonstration combines fetching and rendering a domain with classification. The 2.2-second result includes that complete workflow, rather than measuring only model inference. The comparison is useful for identifying intended workloads, but it does not prove superior phishing detection, universal speed gains, or equivalent performance on local hardware. Its broader tables show task-dependent strengths and weaknesses rather than an across-the-board sweep.
The company also offers engineers to help customers fine-tune Clef. Its planned self-service stack would connect traffic capture, model rollouts, sandboxed scoring, training, and redeployment. The pricing page documents hosted inference rates; it should not be read as a price list for that unfinished training platform.
Perplexity supplies a parallel product signal through its verified Decisions API. Its first-party model and interface are real, but the dossier could not verify an exact launch day or the circulated price. Similar answer-type names do not prove a shared architecture or formal standard.
The allowed choices also define what the application can learn from a response. A routing schema that omits a relevant destination cannot recover it merely by receiving a confident score. Reviewing the schema alongside the model is therefore part of evaluating the deployed decision system, rather than a separate formatting concern.
The Hacker News discussion asks whether this is established classification practice repackaged. That is the strongest novelty objection. The defensible advance is the packaged, multimodal, schema-conditioned scoring path and its deployment options. Independent same-hardware evaluation and task-specific calibration remain missing. For builders, Clef is a shipping alternative worth measuring against the full existing workflow, including parsing, abstention, and the cost of wrong decisions.
Key questions
How does Clef differ from asking a chatbot for JSON?
Has Cloudflare proved Clef’s probabilities are calibrated?
Is Cloudflare’s self-service fine-tuning platform available?
Cite this
APA
Ground Truth. (2026, October 2). Cloudflare launches Clef models that score decisions without writing an answer. Ground Truth. https://groundtruth.day/news/cloudflare-clef-schema-bound-decisions.html
BibTeX
@misc{groundtruth:cloudflare-clef-schema-bound-decisions,
title = {Cloudflare launches Clef models that score decisions without writing an answer},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/news/cloudflare-clef-schema-bound-decisions.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.