開放模型能做安全研究嗎?Cantina 的 apex-flash-1 解決了 60 個保留漏洞任務中的 40 個
Cantina Security, with Yeta Labs, has released apex-flash-1, an open-weights model trained specifically for vulnerability research.It is a reinforcement learning fine-tune of Z.ai’s GLM-5.3-Flash, released on Hugging Face under the MIT license.Is it deployable?
Yes, the MIT weights serve on vLLM, SGLang or Transformers, but BF16 needs roughly 640 GB of GPU memory.What Cantina Built apex-flash-1 has 321.3B total parameters, per its Hugging Face safetensors metadata.The GLM-5.3-Flash base is a Mixture-of-Experts model with 18B active parameters.
Cantina trained it with GRPO using a rank-256 LoRA plus selective full-parameter training.The data covers 150 tasks built from 50 real vulnerability cases.Each case appears in 3 variants: guided whitebox, focused whitebox and focused blackbox.
Authorization, identity and scope flaws make up 72% of cases.Accounting and numerical precision bugs add 18%.Time validation, business rules and SSRF cover the rest.As per the model card on HF, RL rollouts ran inside the Codex agent harness on production-like software and protocol environments.
Benchmark Results Cantina evaluated 60 tasks from 20 held-out vulnerability cases.Each model ran the set once, with costs estimated from provider pricing.apex-flash-1: 40/60 solved (66.7% pass@1), about $2.38 GLM-5.3-Flash (base): 36/60 solved (60.0%), about $4.
56 Claude Opus 5 High: 43/60 solved (71.7%), about $74.68 Opus solved 3 more tasks but cost about 31x more per run.That is roughly $0.06 per solved task for apex-flash-1 versus $1.74 for Opus.These are company-reported numbers on an internal benchmark.
A Worker Model, Not an Orchestrator Cantina positions apex-flash-1 as a worker orchestrated by a larger model.The card lists code reading, tool use, exploit development and verification as target skills.An experimental apex-flash-1-abliterated variant ships with modified refusal behavior.
It was not separately evaluated.Cantina’s rationale is that defenders need capable models they can run and control locally.Interactive Explainer (function(){var f=document.getElementById("mtp-apex-flash-explainer");window.addEventListener("message",function(e){if(f&&e.source===f.contentWindow&&e.
data&&e.data.mtpH){f.style.height=e.data.mtpH+"px";}});})(); How It Compares Featureapex-flash-1Aikido Altar-1Cisco Foundation-Sec-8B-ReasoningGLM-5.3-FlashDeveloperCantina Security + Yeta LabsAikido SecurityCisco Foundation AIZ.aiBase modelGLM-5.3-FlashGLM-5.3 (pruned)Llama 3.
1 8BOwn pretrainingSize321.3B total, BF16328 GB, INT4 (W4A16)8B320B total, 18B activeLicenseMITInherits GLM-5.3 licenseCustom (see NOTICE.
md)MITSecurity methodGRPO RL on 50 real vulnerability casesExpert pruning (REAP) + quantizationInstruction tuning + RLHF on security QAGeneral-purpose basePrimary useAgentic vuln research workerAir-gapped autonomous pentestingSOC triage and threat defenseGeneral coding and agentsHardwareMulti-GPU node (~642 GB BF16 weights)4x H200 with vLLMSingle GPUMulti-GPU nodePublished security result66.
7% pass@1, 60 tasks60.4% recall, 32-CVE internal setCisco-reported security benchmarks60.0% on Cantina’s set Sources: Cantina, Aikido, Cisco, Hugging Face model cards.© Marktechpost Key Takeaways apex-flash-1 is a 321.3B open-weights security model under MIT.
GRPO training on 50 real vulnerability cases produced 150 tasks.It scored 66.7% pass@1 versus 71.7% for Claude Opus 5 High.Its 60-task run cost about $2.38 versus $74.68 for Opus.BF16 needs a multi-GPU node; community 4-bit ports exist.Check out the Technical details.
All credit goes to the researcher of this project.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Can an Open Model Do Security Research?Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks appeared first on MarkTechPost.
Related
相關文章

限時28天!OpenAI承諾沒新功能就重置,網友:只想要Opus
有改進就體驗,沒改進就重置,橫豎不虧。henry 發自 凹非寺 | 公眾號 QbitAI 要我說,OpenAI(SI)已經瘋掉了!專挑在放假的時候整活~ 剛剛,賽博義父、OpenAI Codex負責人Tibo放話:接下來的28天,每天都要二選一: 要麼交付一項對大多數Codex/Work用戶明顯有用的改進,要麼來一次完整的額度重置。

跨越百年的代言:斯凱孚用 AI“復活”已故女星葛麗泰 · 嘉寶拍廣告
作者:遠洋 責編:遠洋 評論: 10 月 5 日消息,“開拍了嗎?”由人工智能生成的好萊塢三十年代影星葛麗泰 · 嘉寶(Greta Garbo)的形象柔聲問道。這是一則看起來頗出人意料的廣告,廣告的投放方是瑞典滾珠軸承製造商斯凱孚(SKF)。

軟銀集團孫正義罕見發出 AI 安全警告,呼籲各國攜手應對威脅
作者:沁滄(實習) 責編:沁滄 評論: 感謝網友 烏蠅哥的左手、不一樣的體驗 的線索投遞!10 月 5 日消息,據彭博社報道,軟銀集團創始人孫正義一直被視為人工智能最堅定的擁護者之一。然而他近日坦言,隨著 AI 能力的突飛猛進,就連他也對伴隨而來的安全風險深感擔憂。

Meta 的 AI 助手 Muse 被曝可深度分析用戶社交關係,併為每位聯繫人建立個人檔案
作者:遠洋 責編:遠洋 評論: 10 月 5 日消息,Meta 的全新個人助手 Muse 已經迅速走紅,數百萬用戶下載了這款人工智能智能體,把它和自己的銀行賬戶、消息軟件或者健康數據進行綁定,交由它代為完成各項事務。不過就在 Muse 在普通消費者群體當中迅速普及之際,來自應用內部的數據,也讓外界得以窺見這款產品整理、向用戶展示信息的運行邏輯。

OpenAI 奧爾特曼:人工智能的巨大效益值得承擔部分風險
作者:沁滄(實習) 責編:沁滄 評論: 感謝網友 不一樣的體驗 的線索投遞!10 月 5 日消息,OpenAI 首席執行官薩姆 · 奧爾特曼(Sam Altman)近日在接受《Decoded》欄目專訪時表示,他的公司與競爭對手 Anthropic 在人工智能監管問題上,仍然存在根本性的世界觀分歧。

9秒刪庫,700個AI失控,豪擲513億買不來安全?
新質動能2026.10.05 08:50 · 來自北京全文4144字00:00 / 12:04AI從工具變成同事,安全就從選項變成前提。文 | 新質動能過去幾年,企業談AI安全,大多數時候指的是內容安全:過濾有害輸出、防止數據洩露、滿足合規備案。