ClawSentry擋惡意技能
Computer Science > Cryptography and Security arXiv:2608.
21101 (cs) [Submitted on 21 Aug 2026] Title:ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents Authors:Kai Wang, Zeming Wei, BiaoJie Zeng, Chang Jin, An Wang, Xiaokun Luan, Zhixiao Lin, Jingjing Qu, Xia Hu, Xingcheng Xu View a PDF of the paper titled ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents, by Kai Wang and 9 other authors View PDF HTML (experimental) Abstract:As large language model (LLM) agents move from conversation to executing code, reading local files, and orchestrating external tools, a single agent hijacked by a malicious third-party skill can cause data exfiltration, privilege escalation, or cascading compromise.
We argue that agentic risk is progressive: it can enter at four loci of the agent control loop--skill admission, invocation-time intent, execution-time effect, and post-action consequence--while a denied dangerous objective can reappear across surface forms, tools, or turns; existing safeguards are typically local to one lifecycle boundary or one call.
Guided by this threat model, we present ClawSentry, an open-source, framework-agnostic security supervision gateway for agent runtimes.
Before a skill package is ever executed, First-use Skill Package Review (FSPR) audits it under a deterministic evidence floor, escalating unresolved cases to bounded read-only agentic review (locus A).
At runtime, a three-tier progressive decision engine--a deterministic L1 layer, a rule-anchored L2 semantic reviewer, and a read-only L3 evidence-seeking agent--spends contextual review only on the residual ambiguity, while a session-level anti-bypass mechanism recognizes tool-switching and rephrased retries (loci B--C); a post-action path feeds high-severity evidence non-retroactively into later review (locus D).
An Agent Harness Protocol (AHP) abstraction applies one policy across Codex, Claude Code, Kimi CLI, and Gemini CLI without modifying agent internals.On SkillInject with Codex/GPT-5.4, contextual ASR falls from 39.55% to 2.61% while contextual TSR moves only from 83.78% to 83.05%.
Across five Work Agents on the full SkillsSafety benchmark, ClawSentry confines ASR to 9.09--15.03% from 33.5--49.7% unprotected, and aggregate TSR on clean skills remains 98.7%.Comments: 35 pages, 14 figures.Code: this https URL Subjects: Cryptography and Security (cs.
CR); Artificial Intelligence (cs.AI) Cite as: arXiv:2608.21101 [cs.CR] (or arXiv:2608.21101v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2608.
21101 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Kai Wang [view email] [v1] Fri, 21 Aug 2026 13:47:51 UTC (798 KB) Full-text links: Access Paper: View a PDF of the paper titled ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents, by Kai Wang and 9 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.
CR < prev | next > new | recent | 2026-08 Change to browse by: cs cs.AI References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...
Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.
ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?
) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.
AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?
) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)
Related
相關文章

剛剛,Codex恢復5小時限額,用戶哀嚎
AI應用風向標(公眾號:ZhidxcomAI) 作者|畢偉豪 編輯|漠影 8月25日報道,剛剛,OpenAI Codex負責人Tibo宣佈,明天開始將會恢復ChatGPT Work和Codex平臺上Plus用戶的5小時用量限制。 據Tibo所說,這一做法早就在計劃之中,目的主要是減輕算力負荷,從而在周使用額度上提供更慷慨的計劃。 這一表述暗含了計算資源不足的意思,但有網友對於Tibo的說法並不買賬,原因是前段時間他親口說過“我們不缺算力”。

中國 AI 為什麼追得這麼快?美媒找到一批關鍵人物
首頁 > 智能時代>人工智能 中國 AI 為什麼追得這麼快?美媒找到一批關鍵人物 2026/8/25 12:46:15 來源:鳳凰科技 作者:於雷 責編:遠洋 評論: 8 月 25 日,據《華爾街日報》報道,中國 AI 模型近年來快速縮小與美國頭部模型的差距,其背後並非短期突然出現的技術突破,而是清華大學等高校長期形成的人才網絡,以及開源技術、人才迴流和更高算力效率等因素共同推動的結果。

Generalist AI融資超13億,8VC領投
機器人前瞻(公眾號:robot_pro) 作者 | 周加琦 編輯 | 漠影 機器人前瞻8月25日報道,近日,據外媒Axios報道,Generalist AI籌集了約2億美元(約13.45億元)的資金。 據消息人士透露,本次融資由8VC領投,數家現有投資者跟投。關於估值暫無具體信息,但確認高於6月份的20億美元水平。目前該公司並未回應,但在聯邦申報文件中披露了本輪融資。 根據Generalist AI向美國SEC提交的Form-D備案文件顯示,本輪計劃募資約2.08億美元(約13.99億元),實際完成融資約1.

人民教育音像數字出版社與小猿達成合作 “中小學課本學習智能體”首落小猿AI學習機
人民教育音像數字出版社與小猿學習機合作,推出首個「中小學課本學習智能體」,搭載於小猿AI學習機,結合權威教材與猿力大模型,提供AI伴學與AI導學雙模式。試點數據顯示,學生語文知識點掌握量提升33.6%,英語提升24%,家長陪讀時間每日減少20分鐘。雙方簽署合作協議,未來將持續推動智能體規模化應用。

WAIC CONNECT | AI出海,不聊虛的!CONNECT帶你拿下馬來西亞真正的AI採購需求
WAIC CONNECT MALAYSIA活動將於2026年9月7日至8日在吉隆坡舉辦,聚焦馬來西亞政府、電信、金融科技與教育等領域的實際AI採購需求,並與華為合作建立當地生態鏈接。活動規劃中國企業短講、一對一精準對接及監管機構與決策者面對面交流,僅限50至60家中國AI企業參與,旨在協助業者直接取得東南亞市場的真實商業訂單。

被智譜封號的會員,有多少是被冤枉的?
鋅消費研究中心2026.08.25 11:33 · 來自北京全文3351字00:00 / 09:02這麼突然被封號了?文 | 鋅消費研究中心 真金白銀辦了智譜會員,結果賬號突然就被封了。 你如果遇到這種事兒,會不會賬號剛一被封,就想原地發瘋? 先別急。智譜之所以會對用戶進行封號,往往源於用戶的登錄設備、IP等出現異常,被平臺判定為多人使用賬號。 而之所以出現這種情況,除了智譜定價越來越貴之外,還源於智譜新、老會員定價不同,讓會員們有了更多套利的空間。