Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of Text

2026年10月2日 01:00
站內 AI 整理稿

Cloudflare has released Clef and Clef-flash, the first models trained by its Workers AI team.They are decision models, not chatbots.Each reads an input state and a schema of typed questions.It returns a probability for every allowed answer, with no free-form text.Both are open-weight under Apache 2.

0 and compatible with TypeSafe AI’s Jev API.Is it deployable?Yes, Both models run today on Workers AI, and the weights are on Hugging Face for self-hosting.What a Decision Model Does An LLM generates tokens one at a time, and its output still needs parsing.

A decision model only answers a fixed set of questions about an input.Clef supports 3 question types: noul: yes/no, returns the probability of yes.choice: picks 1 named option, with per-option probabilities and a confidence value.

score: rates against an ordered rubric, returning a probability-weighted score.On Workers AI, 1 request carries up to 64 questions and up to 4 images.TypeSafe AI launched Jev, its first ‘System One’ model, on September 15, 2026.Open alternatives like Kev-9B and Laya followed.

Clef uses the same System One API.Switching from Jev means changing the endpoint and model name.How Clef Works Clef is post-trained from Qwen3.8-27B, and Clef-flash from Qwen3.5-9B.Both keep the backbone’s vision encoder.Inference has 2 stages.

The backbone first runs a single prefill-only pass over the state and questions.A small transformer, the joint schema head, then reads the final hidden states.It routes evidence to each question, lets fields cross-attend, and scores all options jointly.

A per-question softmax turns logits into probabilities.Training froze both backbones and jointly optimized the routing head with rank-256 low-rank adapters.The loss pairs label-smoothed cross-entropy with a Brier loss for calibration.

A secondary objective, Reinforcement Learning for Calibrated Decisions (RLCD), gives partial credit to adjacent ordinal choices.Interactive Explainer (function(){var f=document.getElementById('mtp-clef-frame');window.addEventListener('message',function(e){if(e.source===f.contentWindow&&e.data&&e.

data.mtpClefH){f.style.height=e.data.mtpClefH+'px';}});})(); Benchmarks: Where Clef Wins and Where It Does Not On Cloudflare’s 10-benchmark shortlist from the Decision Index 0.2.1 suite, a Clef model scored highest on 7.BANKING77 (macro-F1): Clef 94.20 vs Jev 79.74.CLINC150+OOS (macro-F1): Clef 97.

43 vs Jev 89.27.Home appliances (case exact): Clef-flash 97.73 vs Jev 52.27.Jev keeps clear leads elsewhere.The full model card shows Jev ahead on GPQA Diamond (78.3 vs 48.0).It also leads MMLU-Pro (82.7 vs 65.9) and BBH (92.9 vs 73.7).

On TypeSafe’s own workflow evals, Clef beat Jev in 3 of 4 areas, by small margins.Invoice processing was 64.7 vs 61.8, customer service 76.3 vs 76.0, and security incidents 62.9 vs 61.7.Jev leads agent trace observability, 71.6 vs 68.5.

In Cloudflare’s threat intelligence workflow, Clef classified a domain in 2.2 seconds.gpt-oss-120b took 4.7 seconds.All numbers are vendor-reported, with no independent replication yet.

Feature Comparison FeatureClefClef-flashJevKev-9BLayaDeveloperCloudflareCloudflareTypeSafe AIJared PalmerConvai InnovationsSize27B9BNot disclosed9B + 45.4M LoRA421MBackboneQwen3.8-27BQwen3.5-9BNot disclosedQwen3.5-9B-BaseModernBERT-largeWeightsApache 2.0Apache 2.0Hosted APIApache 2.0Apache 2.

0Image inputYesYesNoNoNoContext65,53665,53632K (per Cloudflare)65,536 (8,192 validated)512 (English)Median latency209.3 ms38.8 ms524.1 ms51.4 ms5.8 msHosted price (input)$0.24/M$0.09/M$0.042/MSelf-hostSelf-host Cloudflare’s internal Decision Index run.All 5 implement the System One API.

Sources: Clef docs, Clef-flash docs, Clef card, Jev post, Kev-9B card, Laya card.Deployment and Fine-Tuning Both models are callable through the Workers AI binding (env.AI.run()), the REST API, or AI Gateway.For self-hosting, the model cards list testing on a single H200 with BF16 weights.

Cloudflare also announced a reinforcement learning service for tuning Clef on private data.It starts with Cloudflare’s forward-deployed engineers, with a self-serve platform later.The pipeline combines AI Gateway, Workers AI, Containers and a new Trainer component.

Teams can apply via the design partner form.Key Takeaways Clef (27B) and Clef-flash (9B) are Apache 2.0, Jev-compatible decision models.Median latency: 209.3 ms for Clef, 38.8 ms for Clef-flash, 524.1 ms for Jev.Clef reads text, JSON, images and video within a 64K-token context window.

Jev still leads on knowledge-heavy tests like GPQA Diamond, MMLU-Pro and BBH.An RL fine-tuning service starts with Cloudflare’s forward-deployed engineers.Check out the Model weight, Demo and Technical details.All credit goes to the researcher of this project.

Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?

Connect with us The post Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of Text appeared first on MarkTechPost.

Related

相關文章

鈦媒體生成式AI

領導層洗牌後,Argon能否助谷歌重回前沿?

領導層洗牌後,Argon能否助谷歌重回前沿?影子備忘錄2026.10.02 09:29 · 來自廣東全文5137字00:00 / 13:38谷歌AI的一場豪賭文 | 影子備忘錄2026年的AI賽道,用“瞬息萬變”來形容已經遠遠不夠。如果你每隔兩週沒有關注前沿模型的動態,打開排行榜時大概率會懷疑自己是不是錯過了整整一個季度。就在這樣近乎瘋狂的產品迭代節奏中,谷歌DeepMind於9月30日正式發佈了Gemini 4系列的首款旗艦模型Argon。距離上一代旗艦Gemini 3系列亮相,已經過去了將近十個月。對於一個頭部AI實驗室來說,十個月的沉默幾乎等同於一次豪賭。但在討論Argon之前,有一個更值得回看的背景,然而這場豪賭的籌碼,其實在兩個月前就已經被重新洗過一遍了。一場“靜悄悄的權力交接”2026年8月5日,Alphabet CEO桑達爾·皮查伊通過內部信宣佈了一項令業界震動的架構調整:Google DeepMind聯合創始人兼CEO德米斯·哈薩比斯卸下日常運營管理職責,轉任DeepMind董事長及Alphabet首席科學家。首席技術官科雷·卡武克丘奧盧升任DeepMind高級副總裁,直接向皮查伊彙報。而在這家公司工作了27年的首席科學家傑夫·迪恩,選擇離開谷歌,與老搭檔桑傑·格瑪沃特共同創辦一家獨立的公益企業。這組信息放在一起看,遠不止幾個人換個頭銜那麼簡單。哈薩比斯是DeepMind的靈魂人物,從AlphaGo到AlphaFold,他的個人品牌幾乎等同於DeepMind的公眾形象。把他從日常管理中“解放”出來,轉而專注AGI長期戰略和前沿科學,釋放的信號非常明確:谷歌希望DeepMind從一家“頂級研究機構”變成一臺“高效產品引擎”。卡武克丘奧盧是深度學習的實幹派,在DeepMind工作13年,從組建深度學習團隊到主導WaveNet和DQN等突破,其履歷橫跨研究與工程兩

剛剛
IT之家生成式AI

騰訊 WorkBuddy:Hy4 preview 夜間限免、Hy3 限免延期至 10 月 31 日

首頁 > 智能時代>人工智能 騰訊 WorkBuddy:Hy4 preview 夜間限免、Hy3 限免延期至 10 月 31 日 2026/10/2 9:27:21 來源:IT之家 作者:沁滄(實習) 責編:沁滄 評論: IT之家 10 月 2 日消息,騰訊 WorkBuddy 於 9 月 30 日晚宣佈,決定將 Hy3 模型限免及 Hy4 preview 模型夜間限免均延期至 10 月 31 日。如果用戶已經體驗過 Hy4 preview —— WorkBuddy 在閒時夜間 (23:00 - 次日 8:00) 繼續提供免費體驗 Hy4 preview 的模型額度,截止 10 月 31 日,其餘時間正常消耗積分使用。如果用戶暫未體驗過 Hy4 preview —— 只要在 10 月 10 日 23:59 前首次開啟 Hy4 preview 對話,從首次開啟那天算起,往後 14 天都有每日免費額度。(舉例:10 月 10 日首次開啟,免費額度用到 10 月 24 日)IT之家注意到,官方此前曾發佈公告說明,Hy4 preview 是一個大語言模型,暫不具備多模態能力;當用戶調用 Hy4 preview 執行視頻、圖像等生成任務時,會切換到相應多模態模型完成,因此會按照正常規則產生積分消耗。此外,因免費體驗參與熱度太高,為保障每位用戶都能順暢使用,官方對每日免費額度做了合理分配;當日資源繁忙時會進入排隊,會在頁面提示重置時間恢復。 投訴水文 我要糾錯 下載IT之家APP,簽到賺金幣兌豪禮 相關文章關鍵詞:騰訊 WorkBuddy,Hy4 preview,Hy3,混元騰訊 WorkBuddy 上線微信小程序發佈能力:可使用自然語言生成、上線小程序騰訊混元宣佈“騰訊 Hy 翻譯”App 上線,支持 33 種語言互譯騰訊 QClaw 微信遠程辦公 AI 助手宣佈 12 月 24 日

剛剛
量子位生成式AI

何愷明團隊新作:看貓片就能學會ARC挑戰

何愷明團隊提出NAT-ARC,一種純視覺的ARC解題方案,不依賴語言模型,而是使用ImageNet上的自然圖像進行MAE預訓練,再遷移到抽象格子推理任務。該方法在ARC-1上達到63.4%的單模型pass@2分數,集成後提升至70.2%,逼近專用LLM系統的表現,並證明了視覺預訓練能有效破解抽象推理的scaling瓶頸。

2 小時前
量子位生成式AI

谷歌Gemini 4突然發佈!RSI加持,GPT和Opus都讓讓

。。。 假期第一天,谷歌攜Gemini 4 Argon空降多榜單第一。 拳打Opus 5.5,腳踢GPT-6 Astra。 更誇張的還在後頭,單項任務成本最低可至1.99美元,直接是Astra費用砍半。 最高百萬Token輸出上限,面向編程、金融和法律等複雜工作流,而且劃重點,網絡安全防禦能力超牛掰。 這波等等黨要贏麻了。 不過吧,咱普通用戶現在只可遠觀,暫時還吃不上。

2 小時前
鈦媒體生成式AI

階躍星辰“第一梯隊”,是“自嗨”嗎?

AIX財經2026.10.01 18:22 · 來自福建全文5252字00:00 / 15:27同行各自跑出了主線,階躍的“全棧”能跑通嗎?文 | AIX財經,作者|雷晶,編輯|魏佳沉寂許久的階躍星辰,正試圖擠進大模型第一梯隊。9月20日,它發佈Step 5 Preview,原生支持文本與視覺輸入,重點面向編程、軟件工程、專業知識工作和長程Agent任務。上線初期,登錄並完成首次調用後可獲得30天的Step Plan免費使用權。幾天後,Step Plan的月度套餐一度售罄。

7 小時前
MarkTechPost AI生成式AI

Cohere 發布 Embed 5:與 Voyage 4 Large、Gemini Embedding 2 及 OpenAI 的比較

Cohere 推出全新嵌入模型系列 Embed 5,專注於企業搜尋、RAG 及代理檢索。該系列分為兩個版本:Embed 5 Pro 專注於最高檢索品質,Embed 5 Fast 則針對即時查詢路徑的延遲與成本最佳化。兩者皆支援文字、圖片及文字與圖片混合輸入,涵蓋 100 種以上語言,且可處理多達 128K token。關鍵設計在於 Pro 與 Fast 共享同一嵌入空間,因此可用其中一個進行索引,另一個進行查詢。目前這兩個版本已在 Cohere API、Model Vault、Microsoft Foundry 及 Amazon SageMaker 上正式上線,私有 VPC 或本地部署則可透過 vLLM 提供服務。根據 Cohere 模型文件,API 模型 ID 分別為 embed-v5.0-pro 和 embed-v5.0-fast,兩者輸出維度皆為 2048 或 1536。

8 小時前