Cloudflare's Clef: An Open-Source, Multimodal Answer to Jev Decision Models
A couple of weeks after Typesafe AI’s Jev put “decision models” on the map, Cloudflare has shipped its own — and open-sourced it. Meet Clef and Clef-flash, released October 1 under Apache 2.0, hosted on Workers AI and downloadable on Hugging Face.
First, what’s a decision model?
If you’ve been following the decision-model thread, this is familiar territory. Unlike an LLM, a decision model doesn’t generate text. You hand it an input state and a set of typed questions — is this urgent? which team? a score from 1–5? — and it returns a calibrated probability for every allowed answer. Your code then routes the ticket, triggers an escalation, or defers to a human.
The appeal is determinism and speed: bounded, structured outputs that are cheap, fast, and consistent, without retraining every time you add a category. It’s the piece that lets an agent make a programmatic decision in the middle of a workflow instead of round-tripping a chatty LLM.
Why Clef is interesting
Clef is Jev-API compatible, so existing Jev code works with almost no changes. But Cloudflare used the clean-room rebuild to add real capability:
- Vision encoder — Clef takes images and classifies visual content. Jev is text-only today. This is the headline difference.
- 64k context window — double Jev’s 32k, so you can stuff more input state into a single decision.
- Lower latency — across 43 eval benchmarks, Clef’s models beat the competition on latency, and Clef-flash is startlingly fast while staying competitive on quality.
- Leading the index — Clef currently tops the Jev Decision Index, beating Jev in 3 of 4 areas.
The name is a nice touch: in music a clef assigns pitch to the lines of a staff — it defines the domain before the notes follow. A decision model does the same for context and the actions that follow it. (And “Clf” tips its hat to Cloudflare.)
A concrete win: threat intelligence
Cloudflare’s own dogfooding is the clearest illustration. Their Threat Intelligence team uses Clef (with Browser Run) to classify domains: give it a URL and it returns something like 95% fashion, 85% ecommerce, <1% phishing. That fetch-render-classify loop took Clef 2.2s; their fastest general LLM, gpt-oss-120b, took 4.7s in the same workflow and returned only two classifications. A 2x latency saving with richer output — repeated across millions of domains, that’s a material operational edge.
What’s next: RL fine-tuning
Alongside the models, Cloudflare debuted a reinforcement-learning platform that lets you fine-tune Clef on your own data. That’s the part worth watching: a hosted, Jev-compatible decision model you can cheaply specialize to your taxonomy — your ticket categories, your fraud signals, your invoice types — without standing up training infrastructure.
Who should care
If you route tickets, classify invoices, categorize content, or gate tool calls all day, Clef is a near drop-in upgrade worth benchmarking. It’s open (Apache 2.0, run it locally), multimodal, faster, and backed by a managed endpoint plus an upcoming fine-tuning path. The decision-model space got crowded fast — but Cloudflare putting a vision-capable, index-leading model into the open is the kind of move that resets the baseline.
Models: huggingface.co/Cloudflare/clef · clef-flash · Announcement: blog.cloudflare.com/clef-decision-models