ClawSentry擋惡意技能

2026年8月25日 00:00
站內 AI 整理稿

Computer Science > Cryptography and Security arXiv:2608.

21101 (cs) [Submitted on 21 Aug 2026] Title:ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents Authors:Kai Wang, Zeming Wei, BiaoJie Zeng, Chang Jin, An Wang, Xiaokun Luan, Zhixiao Lin, Jingjing Qu, Xia Hu, Xingcheng Xu View a PDF of the paper titled ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents, by Kai Wang and 9 other authors View PDF HTML (experimental) Abstract:As large language model (LLM) agents move from conversation to executing code, reading local files, and orchestrating external tools, a single agent hijacked by a malicious third-party skill can cause data exfiltration, privilege escalation, or cascading compromise.

We argue that agentic risk is progressive: it can enter at four loci of the agent control loop--skill admission, invocation-time intent, execution-time effect, and post-action consequence--while a denied dangerous objective can reappear across surface forms, tools, or turns; existing safeguards are typically local to one lifecycle boundary or one call.

Guided by this threat model, we present ClawSentry, an open-source, framework-agnostic security supervision gateway for agent runtimes.

Before a skill package is ever executed, First-use Skill Package Review (FSPR) audits it under a deterministic evidence floor, escalating unresolved cases to bounded read-only agentic review (locus A).

At runtime, a three-tier progressive decision engine--a deterministic L1 layer, a rule-anchored L2 semantic reviewer, and a read-only L3 evidence-seeking agent--spends contextual review only on the residual ambiguity, while a session-level anti-bypass mechanism recognizes tool-switching and rephrased retries (loci B--C); a post-action path feeds high-severity evidence non-retroactively into later review (locus D).

An Agent Harness Protocol (AHP) abstraction applies one policy across Codex, Claude Code, Kimi CLI, and Gemini CLI without modifying agent internals.On SkillInject with Codex/GPT-5.4, contextual ASR falls from 39.55% to 2.61% while contextual TSR moves only from 83.78% to 83.05%.

Across five Work Agents on the full SkillsSafety benchmark, ClawSentry confines ASR to 9.09--15.03% from 33.5--49.7% unprotected, and aggregate TSR on clean skills remains 98.7%.Comments: 35 pages, 14 figures.Code: this https URL Subjects: Cryptography and Security (cs.

CR); Artificial Intelligence (cs.AI) Cite as: arXiv:2608.21101 [cs.CR] (or arXiv:2608.21101v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2608.

21101 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Kai Wang [view email] [v1] Fri, 21 Aug 2026 13:47:51 UTC (798 KB) Full-text links: Access Paper: View a PDF of the paper titled ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents, by Kai Wang and 9 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.

CR < prev | next > new | recent | 2026-08 Change to browse by: cs cs.AI References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...

Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.

ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?

) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.

AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?

) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)

Related

相關文章

鈦媒體生成式AI

被智譜封號的會員,有多少是被冤枉的?

鋅消費研究中心2026.08.25 11:33 · 來自北京全文3351字00:00 / 09:02這麼突然被封號了?文 | 鋅消費研究中心 真金白銀辦了智譜會員,結果賬號突然就被封了。 你如果遇到這種事兒,會不會賬號剛一被封,就想原地發瘋? 先別急。智譜之所以會對用戶進行封號,往往源於用戶的登錄設備、IP等出現異常,被平臺判定為多人使用賬號。 而之所以出現這種情況,除了智譜定價越來越貴之外,還源於智譜新、老會員定價不同,讓會員們有了更多套利的空間。

剛剛
IT之家生成式AI

中消協發佈消費提示:使用人工智能服務需謹防誤導

作者:遠洋 責編:遠洋 評論: 8 月 25 日消息,中國消費者協會今天發佈人工智能服務消費提示,提醒廣大消費者理性認識通用生成式人工智能的信息輔助和參考屬性,對涉及價格費用、合同條款、售後服務、法律責任等重要消費信息,不宜僅依據人工智能生成內容作出決定,應通過經營者官方渠道、合同文本、監管部門及其他權威渠道進行核實確認。

剛剛
量子位生成式AI

VC開始靠AI預測未來了

DigClaw的預測框架Rhizome在FutureX評測中,以同一套系統跨三個不同基礎模型取得前三名成績,驗證其預測能力可獨立於模型權重。該系統將搜索、因果推理與概率推斷解耦,並透過軌跡記錄與校準累積數據資產,試圖解決大型語言模型不擅長因果預測的根本問題。這項技術由AI原生創投Newborn Ventures支持,目標是將投資背後的趨勢預測系統化,應用於投資決策與產業風險評估。

剛剛
IT之家生成式AI

全球最大生活指南百科 WikiHow 指控 OpenAI 爬取超 1.1 萬篇教程文章訓練 AI

作者:故淵 責編:故淵 評論: 感謝網友 愚公騎馬 的線索投遞!8 月 25 日消息,路透社今天(8 月 25 日)發佈博文,報道稱 WikiHow 於 8 月 21 日向美國曼哈頓聯邦法院提交訴訟,指控 OpenAI 其未經許可抓取超過 11,000 篇教程文章,用於訓練 ChatGPT 及 GPT 大型語言模型,並侵犯至少 1,200 項已註冊版權。

剛剛
鈦媒體生成式AI

Edge AI Daily 早報(8月25日)

Edge AI Daily2026.08.25 08:05 · 來自北京全文4868字00:00 / 13:43微軟Skala 1.1改寫計算化學底層規則;Coherent用SiC襯底提升AI散熱25%。AI倫理調查顯示ChatGPT等模型高頻鏈接反墮胎組織。

剛剛