Open to workNew YorkGet in touch

Decision Models

Clef decision models ship Apache 2.0 weights, a 38.8 ms flash tier and RL fine-tuning

Cloudflare open-sourced two Jev-compatible decision models on Qwen backbones and opened an RL fine-tuning service that, for now, runs through its forward-deployed engineers.

Published
October 1, 2026
Read
7 min
Author
Samir Sengupta
Cloudflare's Clef and Clef-flash decision models compared with Jev on latency and benchmark scores
Note 077 / 081Daily note · Written from 2 sources

the short version

  • Clef and Clef-flash are Apache 2.0 open weights built on frozen Qwen3.8-27B and Qwen3.5-9B backbones with rank-256 adapters and a routing head.
  • On Cloudflare's own latency run across 43 benchmarks, Clef-flash's 38.8 ms median is more than 13x faster than Jev's 524.1 ms, though Laya is faster still at 5.8 ms.
  • Clef beats Jev on 8 of Cloudflare's 10 shortlisted benchmarks, but Jev still wins When2Call and BRIGHT, and Clef-flash drops to 66.77 on CLINC150+OOS.
  • Cloudflare pitches fine-tuning for narrow workloads backed by years of labelled decisions. Today it runs through the FDE team, and the self-serve platform has no date.

Cloudflare released two Clef decision models, Clef and Clef-flash, on 1 October 2026. They ship as open weights on Hugging Face under an Apache 2.0 license and are hosted on Workers AI behind a Jev-compatible API, alongside a reinforcement learning fine-tuning service. Clef is built on a frozen Qwen3.8-27B and Clef-flash on a frozen Qwen3.5-9B. Cloudflare reports median latency of 209.3 ms and 38.8 ms against 524.1 ms for Typesafe AI's Jev, and says Clef currently leads the Jev Decision Index, beating Jev on 8 of its 10 shortlisted benchmarks. Fine-tuning is aimed at specific workloads such as Cloudflare's own Trust and Safety, Support and bot classification, and for now it runs through Cloudflare's forward-deployed engineer (FDE) team rather than as a self-serve product.

A decision model takes input state and a set of typed questions and returns bounded answers with probabilities. Code can use those answers to route a ticket, escalate, or defer to a human. The post names two concrete differences between Clef and Jev: Clef has a vision encoder, so it can classify images, where Jev only does text today, and it has a 64k context window against Jev's 32k. Cloudflare says it does not read, store, or train on hosted requests or responses unless a customer opts into fine-tuning.

How Clef scores against Jev-style models

Cloudflare shortlisted ten benchmarks it considers important under the Jev Decision Index. It scored Clef, Clef-flash, Jev, DiffusionGemma Jev, Kev 9B and Laya on them. These are Cloudflare's own runs, also published on its live benchmark demo site, not an independent evaluation.

  • Clef leads on ToolRet nDCG@10 (69.19 vs Jev 65.28), BANKING77 macro-F1 (94.20 vs 79.74), CLINC150+OOS macro-F1 (97.43 vs 89.27) and Amazon ESCI macro-F1 (57.48 vs 55.21).
  • Clef-flash leads on BFCL case exact (98.76), API-Bank accuracy (93.11) and Home appliances case exact (97.73, against Jev's 52.27).
  • Jev still wins When2Call accuracy (80.97 vs Clef 72.37) and BRIGHT nDCG@10 (47.52 vs 45.91).
  • DiffusionGemma Jev wins PhishNChips accuracy at 85.35, ahead of Clef at 79.60 and Jev at 62.55.
  • Clef-flash falls to 66.77 on CLINC150+OOS, well below Clef's 97.43 and Jev's 89.27, so the smaller model is not a uniform substitute.

Cloudflare also ran Typesafe's own workflow eval suite. Clef beat Jev on invoice processing (64.7 vs 61.8) and security incidents (62.9 vs 61.7, where Clef-flash tied Jev at 61.7). Clef-flash edged Jev on customer service (77 vs 76.0), and Jev kept agent trace observability (71.6 vs Clef-flash 69.8 and Clef 68.5). Every one of those four gaps between Jev and the best Clef variant is 2.9 points or less. A team's own labelled sample will tell it more than the headline that Clef wins 3 of 4.

Latency comes from skipping token-by-token decoding

Clef runs the Qwen backbone as a prefill-only pass, then scores the valid schema choices in parallel. The decision step is non-autoregressive, so no intermediate text is generated. A two-stage attention routing process lets each option pull in prompt context and lets fields cross-attend to each other before schema-bound scoring. Across 43 eval benchmarks, Cloudflare reports median latency of 209.3 ms for Clef, 38.8 ms for Clef-flash, 524.1 ms for Jev, 84.4 ms for DiffusionGemma Jev, 51.4 ms for Kev-9B and 5.8 ms for Laya. By that table, Clef-flash beats every model except Laya on median latency, but the larger Clef is slower than Kev-9B and DiffusionGemma Jev.

Clef's p95 of 238.6 ms sits close to its 209.3 ms median, while Clef-flash's p95 rises to 122.4 ms from a 38.8 ms median. Even so, Clef-flash has the lowest p95 in the table, against 536.0 ms for Jev and 222.5 ms for Laya. In a domain classification workflow with Browser Run, Cloudflare's Threat Intelligence team reports Clef took 2.2 s to fetch, render and classify a site. gpt-oss-120b took 4.7 s in the same workflow and returned only two classifications. Cloudflare says Laya is faster but trades away quality on its benchmarks.

Calling the hosted Clef decision models

The hosted model sits at @cf/cloudflare/clef on Workers AI. A request carries a state string plus named questions, typed in the post's example as choice, score, and a field type written as noul for a yes-or-no question. Cloudflare says Clef produces strictly typed outputs and is fully API-compatible with Jev, so the swap should be easy. The post does not show the endpoint or model id for Clef-flash, and the noul type name should be checked against the developer documentation before copying the example.

bash
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef \
  -X POST \
  -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
  -d '{
    "model": "clef",
    "state": "Checkout has been failing for every customer for the last hour.",
    "questions": {
      "urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
      "team": {
        "type": "choice",
        "instructions": "Which team should handle this request?",
        "criteria": {
          "billing": "Payments, invoices, and refunds",
          "technical": "Outages, errors, and configuration",
          "sales": "Plans and upgrades"
        }
      },
      "severity": {
        "type": "score",
        "instructions": "How severe is the customer impact?",
        "criteria": ["No impact", "Minor", "Major", "Critical"]
      }
    }
  }'

Cloudflare positions Clef as the precision model and Clef-flash for latency-critical decisions. It says GPUs at the edge cut network latency, so the call can sit in an agent's hot path next to a Workers AI LLM that takes the action. The blog post does not state pricing. One Hacker News commenter put Clef at $0.24 per million input tokens, roughly 6x Jev, and Clef-flash at $0.09. Treat those figures as unverified until they are checked against Cloudflare's pricing page.

When to fine-tune Clef instead of calling it

Cloudflare's stated case for fine-tuning is specific workloads with a long history of labelled decisions. Its examples are Trust and Safety submissions, Support triage, and good-bot versus bad-bot calls in its Bot products, drawing on more than 15 years of network data. Cloudflare says such a model can be more accurate and faster than generic Clef, but it publishes no numbers for a fine-tuned model. The post is explicit about the trade: a fine-tuned model may give up general-purpose performance for higher accuracy in one domain. If inputs and categories shift often, the hosted general model fits better, since decision models are meant to handle new categories without retraining.

The base Clef recipe jointly trains a routing head and rank-256 low-rank adapters over the frozen backbone. It uses label-smoothed cross-entropy plus a Brier loss for calibration, on synthetic data that permutes field orders, prompts and schema structures. Reinforcement Learning for Calibrated Decisions (RLCD) serves as a secondary optimization target: it gives partial credit for adjacent ordinal choices, rewards fully precise record outputs, and applies a reference penalty against distribution shift. The customer fine-tuning pipeline is assembled from Cloudflare platform primitives, one of them new:

  1. AI Gateway captures AI traffic and automatically builds a dataset of requests for the use case.
  2. Workers AI generates rollouts against the base Clef model.
  3. Containers act as the RL sandbox for scoring and replaying agent actions.
  4. Trainer, a new component, updates the weights of the fine-tuned Clef model.
  5. Workers AI with Bring Your Own Model (Cog), work that has progressed since Cloudflare acquired Replicate, redeploys the result.

Cloudflare describes several of these pieces as work in progress. Today, Clef RL fine-tuning is a hands-on engagement with the FDE team, and the self-serve platform comes later, with no date given. Cloudflare also invites existing customers of these products who have specific use cases to act as design partners.

What Hacker News commenters are pushing back on

The top-ranked thread asked how so many Jev alternatives appeared within weeks. Replies largely agreed that the technique is not new. One commenter said any pretrained LLM can be adapted to score a structured set of options, and several said Jev's contribution was the API and product concept. Cloudflare's own account fits that: its first prototype adapted DiffusionGemma by exposing logprobs, before Clef moved to a Qwen backbone.

  • Narrow tasks may not need Clef at all. One commenter said a BERT-based classifier trained on a laptop in about an hour, given good data, would answer faster than the round trip to Clef or Jev and use under 1 GB of memory.
  • Calibration is contested. One commenter said opinions conflict on how well calibrated each model is, that Jev seems best, and that most users care about accuracy over confidence. Cloudflare describes its Brier loss and RLCD but publishes no calibration metric.
  • Price drew comment. One reply guessed the $0.24 per million input figure reflects Cloudflare's internal need to show the model images of emails and webpages to detect phishing despite obfuscated HTML, which is speculation the post does not support. Another commenter called the pricing more expensive but worth trying for the vision encoder.

What is still unknown about Clef

The post omits several things an engineer would need. It gives no calibration numbers despite training explicitly for calibration, no before-and-after results from any fine-tuned Clef, and no hardware requirements for running the 27B-backbone model locally. It also gives no date for the self-serve RL platform and does not state the hardware or network conditions behind its latency figures. All accuracy and latency figures come from Cloudflare's own runs on benchmarks it chose. Before swapping out Jev, run both against a labelled sample of your own decisions, watch the When2Call and CLINC150+OOS gaps if your workload resembles them, and confirm pricing directly.

Questions this raises

Are Clef decision models faster than Jev?

By Cloudflare's numbers, yes. Across 43 eval benchmarks Clef's median latency is 209.3 ms and Clef-flash's is 38.8 ms, against 524.1 ms for Jev. Clef is slower than Kev-9B and DiffusionGemma Jev, and Laya is faster than both Clef variants at 5.8 ms.

Can I fine-tune Clef on my own data?

Cloudflare offers a reinforcement learning fine-tuning service, but for now it runs through its forward-deployed engineer team rather than as a self-serve product. It targets specific workloads with a long history of labelled decisions, and a fine-tuned model may give up general-purpose performance for higher accuracy in one domain.

Is Clef a drop-in replacement for Jev?

Cloudflare says Clef is fully API-compatible with Jev and produces strictly typed outputs, and the hosted model sits at @cf/cloudflare/clef on Workers AI. Jev still wins some benchmarks, such as When2Call and BRIGHT, and Clef-flash drops to 66.77 on CLINC150+OOS, so testing on your own labelled sample is advised.

These daily notes are drafted by a model I run and operate myself - the same kind of pipeline this site is about - from sources published in the previous 24 hours, and every one lists what it read. The longer essays, the talks and the preprint are mine, written by hand.

More notes

Building something on this?

I ship production LLM, RAG and agentic systems for a living - the infrastructure behind the things these notes are about. Open to roles, contract work and research collaboration.