SALLIE免生成大模型越獄檢測

2026年8月12日 00:00
站內 AI 整理稿

Computer Science > Cryptography and Security arXiv:2604.

06247 (cs) [Submitted on 6 Apr 2026 (v1), last revised 10 Aug 2026 (this version, v2)] Title:SALLIE: Generation-Free Hidden-State Detection of Jailbreaks and Prompt Injections Across Text and Vision Authors:Guy Azov, Ofer Rivlin, Guy Shtar View a PDF of the paper titled SALLIE: Generation-Free Hidden-State Detection of Jailbreaks and Prompt Injections Across Text and Vision, by Guy Azov and 2 other authors View PDF HTML (experimental) Abstract:Large Language Models (LLMs) and Vision-Language Models (VLMs) are vulnerable to jailbreaks and prompt injections delivered through text or images.

Existing defenses often narrow threat coverage or add inference cost through input transformations, repeated model queries, generated responses, or threat-specific detectors.

We present SALLIE (Single-Pass Activation Lookup for Layerwise Input Evaluation), a white-box, generation-free detector with a common architecture for textual and visual jailbreaks and prompt injections.

Using one backbone forward pass, SALLIE extracts last-token residual-stream representations, applies layer-wise k-nearest-neighbor probes, and averages their scores across a contiguous layer range, without modifying the backbone or generating a response.

To distinguish reusable detector configurations from deployment-specific tuning, we compare shared, threshold-only, and fully calibrated regimes across backbones and modalities.We evaluate SALLIE on compact open-weight VLMs (Gemma-3-4b-it, Phi-3.5-vision-instruct, SmolVLM2-2.

2B-Instruct) using heterogeneous held-out test data spanning both threat families and modalities.Fully calibrated SALLIE-Phi achieves an aggregate F1 of 0.87 and higher aggregate balanced accuracy than RCS-KCD on all three backbones under a matched-calibration protocol.On Phi-3.

5 visual inputs, threshold-only calibration yields an F1 of 0.99 with zero observed false positives, versus 0.54 for zero-shot Gemini-2.5-Flash-Lite.On Gemma-4-E2B-it, a backbone held out from representation-hyperparameter selection, the same procedure raises F1 from 0.62 to 0.79 on text and 0.

68 to 0.98 on visual inputs.Representation hyperparameters can thus transfer across some backbones, while operating points remain backbone-, modality-, and distribution-dependent.We further report FPR/FNR tradeoffs, latency, and representation ablations.Comments: 17 pages, 5 figures, 17 tables.

Preprint under review.

v2: substantially revised - new title expansion, reworked method presentation, added calibration-regime analysis (shared / threshold-only / fully calibrated) and matched-protocol comparison with RCS-KCD; technical appendices A-H now included Subjects: Cryptography and Security (cs.

CR); Artificial Intelligence (cs.AI) ACM classes: I.2.7; I.2.10; K.6.5; I.2.6 Cite as: arXiv:2604.06247 [cs.CR] (or arXiv:2604.06247v2 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2604.

06247 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Ofer Rivlin [view email] [v1] Mon, 6 Apr 2026 16:29:05 UTC (566 KB) [v2] Mon, 10 Aug 2026 09:29:00 UTC (605 KB) Full-text links: Access Paper: View a PDF of the paper titled SALLIE: Generation-Free Hidden-State Detection of Jailbreaks and Prompt Injections Across Text and Vision, by Guy Azov and 2 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.

CR < prev | next > new | recent | 2026-04 Change to browse by: cs cs.AI References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...

Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.

ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?

) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.

AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?

) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)

Related

相關文章

量子位生成式AI

阿里視頻大模型Wan3.0正式上線,行業評價“穩定、真實、有質感”

阿里巴巴影片生成大模型Wan3.0正式上線,單次可生成30秒影片,並首次支援doc、xls、ppt、pdf、md等文檔輸入。企業用戶普遍評價其「穩定、真實、有質感」,能穩定保持角色與場景一致性,並已進入短劇、影視、廣告等生產流程。即日起可於阿里雲百鍊、千問等平台體驗,標準版並推出限時7折優惠。

剛剛
IT之家生成式AI

阿里雲視頻生成模型 Wan3.0 正式上線,支持單次生成 30 秒視頻、文檔輸入

作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,阿里雲消息,今天,視頻生成模型 Wan3.0 正式上線。官方稱,Wan3.0 在生成時長、萬能創作、全能參考以及真實世界還原等維度全面升級,單次可生成 30 秒視頻,並首次支持 doc、xls、ppt、pdf、md 等文檔格式輸入,力求準確還原真實世界。

剛剛