Mistral發佈開源安全小模型
重點摘要
法國AI公司Mistral推出開源安全模型Shieldstral,僅有30億參數,但在標準文字安全基準上與規模約七倍大的OpenAI模型表現相當,並在圖文分類任務創新高。該模型允許營運者以自然語言問題自訂審查規則,並回傳單一標記來計算安全分數。Shieldstral以Apache 2.0授權開源,採用合成資料訓練,在適應性基準上略遜於生成較長推理序列的模型,但更具成本效益。
Shieldstral, a 3-billion-parameter model from French AI company Mistral, matches models three times its size on standard text safety benchmarks, according to the paper. Mistral says the model also sets a new high score for joint text and image classification. Runtime rules let operators tailor safety checks Many guardrail models sort content using fixed taxonomies. The paper's authors, including Mistral co-founder Guillaume Lample, point to two problems with this approach: public safety datasets group risks too differently to support one common taxonomy, and the same rules don't fit every use case. Content suitable for a cybersecurity tool could be harmful on a mental health platform.Ad Operators write the review criteria in plain language. Shieldstral returns one token, which produces a safety score between zero and one. | Image: Mistral Operators tell Shieldstral what to check with plain-language questions such as "Does this content promote violence?" The model answers only "yes" or "no," and the system uses the probability of each response to calculate a safety score between zero and one.AdDEC_D_Incontent-1 Synthetic data helps the model handle new rules The researchers combined about 54.1 million examples covering safety, harmful content, and manipulation attempts into one format. They applied strict standards to targeted manipulation, moderate standards to general safety data, and lenient standards to response quality. Each training example includes task instructions, a specific yes or no question, and the content being reviewed. The answer is a single token. | Image: Mistral To teach Shieldstral finer distinctions, the team used another language model to rewrite safe text into unsafe variants. Each example also included a similar but different category that had to be rejected, which trained the model to separate closely related rules rather than make only a broad safe or unsafe judgment.Ad The authors created the adaptability test categories separately from the training set, using different names and levels of detail. None of the fine-grained test categories directly matches a training category, though 10 of the 12 broader classes have rough counterparts, they say. The 3B model ties one nearly seven times larger Across the combined text benchmarks, Shieldstral posts an F1 score of 84.9 percent. F1 combines precision and recall into one metric, with 100 percent representing a perfect score. That result ties OpenAI's GPT-OSS-Safeguard-20B, which is about seven times larger, and beats Qwen3Guard-8B at 84.0 percent, Nemotron-3.5-Safety-4B at 83.3 percent, and LlamaGuard-4-12B at 69.1 percent.AdDEC_D_Incontent-2 On images and image-text combinations, Shieldstral scores 83.8 percent, ahead of OmniGuard-7B at 77.6 percent and LlavaGuard-7B at 71.6 percent.Ad Across the text benchmarks, Shieldstral ties an OpenAI model roughly seven times its size. | Image: Mistral GPT-OSS-Safeguard-20B leads the adaptability benchmark with 94.1 percent, compared with Shieldstral's 91.3 percent. This test uses rules that differ from the training categories or are entirely new. The authors still consider Shieldstral more practical than GPT-OSS-Safeguard-20B and Nemotron-3.5-Safety, since both generate long intermediate reasoning sequences that raise compute costs, while Shieldstral returns a single word. On policies absent from training, Shieldstral trails the two models that generate intermediate reasoning before answering. | Image: Mistral Shieldstral is based on Mistral's Ministral-3B with the Pixtral vision encoder. In a validation test with fine-grained categories, synthetic category data raised the F1 score by 23.3 percentage points, which the researchers say was the main driver of the model's ability to adapt to new rules. Shieldstral is available as an open-weight model under the Apache 2.0 license. Custom rules could limit costly filtering mistakes Safety classifiers sit on either side of the main language model, screening prompts before processing and responses before they reach users. Operators can update these rules without retraining the main model. But because every request passes through the classifier, its size, speed, and cost add up quickly. Anthropic's Claude Fable 5 showed how poorly tuned filters can affect real use. Artificial Analysis found that the system automatically routed eight to nine percent of tasks to a weaker model. One medical physicist called Fable 5 unusable because his work often includes the word "nuclear." Other users reported that the system flagged MRI analysis as bioterrorism. Anthropic tightened the filter after locating a safety issue and says it has since blocked harmless coding tasks more often. Shieldstral gives operators more control over that tradeoff. They can write screening criteria at runtime and tailor the filter to a specific app instead of adopting someone else's categories. These classifiers already play a growing role across the industry. OpenAI uses them for automatic age detection in ChatGPT and routes emotional requests through a safety filter to stricter models. Claude Code uses a classifier to block external scripts, production deployments, and force pushes. Anthropic's Fable 5 review also requires the company to store inputs and outputs for up to 30 days, or up to two years after rule violations. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Arxiv | Mistral BETA-TEST × wpDiscuzInsert BETA-TEST × wpDiscuzInsert
Related
相關文章

百度入局 AI 辦公:dodo 併入百度搭子,全面整合辦公 Agent 業務
百度宣布將內部辦公智能體 dodo 併入百度搭子,整合研發資源,實現內外部辦公智能體產品統一。百度搭子為專業辦公智能體平台,已接入百度多項產品能力,並推出企業版,合併後將集中資源全面進軍 AI 辦公賽道。

耗時41分鐘,千問辦公押注了怎樣的Agent未來?
千問辦公在生成具身智能行業週報的測試中耗時41分鐘,速度雖慢但強調深度判斷與信源核實,與其他追求效率的AI辦公產品形成對比。文章探討了Token作為AI時代商業語言,以及企業可能推動Agent走向更「重」的未來,以建立用戶信任。三款產品代表不同路線,千問辦公選擇了更注重判斷過程的路徑。

辦公智能體平臺“Work”大作戰
騰訊WorkBuddy崛起帶動辦公智能體平台賽道競爭加劇,阿里、字節等大廠加速整合旗下產品迎戰,市場規模快速成長。各大廠商在商業化模式上持續摸索,飛書、滴普等業者表現不一,未來競爭將更為激烈。

DeepSeek:計劃近期整體上調 API 服務的定價,預計漲幅較大
DeepSeek 宣布近期將整體上調 API 服務定價,預估漲幅較大,直接影響開發者與企業用戶的使用成本。目前收費依輸入與輸出 Token 數量及緩存是否命中分別計算,未來漲價後不同使用情境的費用增幅可能有所差異,開發者需重新評估應用成本。

前安克高管做智能房車,獲元禾、厚雪等超2億融資,首款產品2027年初量產|硬氪首發
戶外出行硬科技企業松鼠動力完成超2億元A輪融資,投資方包括元禾璞華、厚雪資本等,資金將用於整車驗證與產能擴建。該公司首款智能增程拖掛房車Evotrex-PG5已於CES 2026發表,售價12萬美元起,預訂量超預期,預計2027年開始大批量量產交付。

飛書併入豆包、千問辦公整合,大廠AI大戰從“賽馬”到“合兵”?
字節跳動近日發佈內部信,宣布對旗下豆包、飛書與火山引擎相關業務進行整合。其中,飛書產品團隊與豆包產品團隊正式合併,組成新的豆包產品團隊,由原本的豆包負責人趙祺統一管理。這是字節跳動自2021年設立六大業務單元以來,針對B端業務所進行的一次關鍵架構調整。字節跳動在內部信中說明,AI正深刻改變企業生產力與個人工作方式,這項整合旨在打通辦公與生產力產品的研發體系,進一步強化服務企業客戶的能力。 這波調整並非單純的團隊合併,而是一次從產品、技術到底層商業化體系的全面重組。