Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills
Anthropic has published a new plugin evals workflow for Claude Code.The claude plugin eval command runs a plugin against realistic prompts, grades what Claude produced, and compares the result with a run where the plugin is not loaded.
It answers 3 questions plugin developers could not previously measure: does the skill trigger, does it survive an edit or a new model, and does it beat a bare model.Deployable: Yes.It runs on Claude Code v2.1.269 or later against any directory with a plugin.json or .claude-plugin/plugin.
json manifest, or a skills-directory plugin.Every eval run and judge grader is a real model call billed to your plan or API account.What a case looks like An eval suite lives in an evals/ directory inside the plugin.Each case is a subdirectory holding a prompt.md and a graders/ folder.
The prompt body goes to Claude exactly as written, and @path mentions are not expanded.Frontmatter on prompt.md can set maxturns (default 10), timeoutseconds (default 300), model, tags, and allowedtools.
Graders are markdown files whose frontmatter sets a type, an optional weight, and an optional arm.There are 6 types.Four cost nothing because they are computed from the transcript and the files on disk: regex, toolused, toolorder, and fileexists.
Two call a judge model and add to the bill: llm, which scores the reply against prose criteria you write, and baseline, which compares it against a reference answer.
claude plugin eval init reads the plugin, asks what a good result looks like, proposes cases and graders, tries them, and writes the files.In CI, --bare writes a blank template instead.
The number that matters is Δ By default every case runs twice: a with-arm where the plugin is loaded and a without-arm where it is not.Their difference, Δ, is what the plugin contributed.If a case scores 1.0 in both arms, the plugin is not why it passed.
The docs example output shows a single case at WITH 1.00, W/OUT 0.33, Δ +0.67 across 6 runs, costing an estimated $0.41 and taking 74 seconds.A grader marked with-only, typically toolused: Skill, is reported as an indicator and excluded from the score, since the without-arm has no skill to fire.
Anthropic calls out the most common first finding: a Δ near zero with the toolused: Skill grader failing, which means Claude is not choosing the skill on natural phrasing.That is the defect claude plugin validate cannot see, because it checks manifest syntax and schema rather than behavior.
Results land under evals/results//report.html with per-grader verdicts and judge votes.Where the account supports it, the report is also published to claude.ai unless --no-publish is set.
Cost and CI A suite makes roughly cases × runs × arms agent runs, plus 3 short judge calls per llm or baseline grader per run, and results vary between runs.The documented CI invocation is: Copy CodeCopiedUse a different Browserclaude plugin eval .\ --trust-plugin \ --json results.
json \ --threshold 0.8 \ --model claude-sonnet-5 \ --judge-model claude-haiku-4-5 \ --no-publish \ --max-cost-usd 20 The runner needs a Claude Code install and credentials such as ANTHROPICAPIKEY.Without --trust-plugin, an untrusted checkout is refused with exit 1 when there is no terminal.
Report problems never change the exit code, and --json suppresses progress output.Interactive explainer (function(){var f=document.getElementById("mtp-plugin-evals-frame");window.addEventListener("message",function(e){if(e.data&&e.data.type==="mtp-plugin-evals-resize"&&f){f.style.height=e.data.
height+"px";}});})(); Key Takeaways claude plugin eval scores realistic prompts with 6 grader types; 4 are free, llm and baseline bill a judge model.Every case runs with and without the plugin by default; Δ is the only number that proves the plugin did the work.
A Δ near zero with a failing tool_used: Skill grader means the skill never triggers on natural phrasing.--threshold, --max-cost-usd, and --trust-plugin turn it into a CI gate; usage-limit errors can fake a regression.Requires Claude Code v2.1.
269+; claude plugin eval init writes the first suite for you.Check out the Technical details.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills appeared first on MarkTechPost.
Related
相關文章

無問芯穹與華環電子簽署戰略合作,共同探索國產異構算力AI基礎設施新方向
無問芯穹與華環電子簽署戰略合作協議,雙方將結合各自在AI軟體平台、網路通信與硬體研發的優勢,共同探索國產異構算力基礎設施的協同方案。此次合作聚焦於智算中心解決方案及「Token工廠」新模式,目標是推動計算、網路與AI原生基礎設施深度融合,為AI規模化應用提供高效穩定的支撐。
Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs
Jina AI, part of Elastic, has released jina-ocr-v1, an end-to-end visual document parser. It takes PDFs, scans, tables, charts or invoices and returns clean Markdown in 1 pass. The model has 3.

優步全球範圍裁員 10%,被裁員工稱 AI 已大舉滲透日常工作
作者:清源 責編:清源 評論: 9 月 18 日消息,據《商業內幕》今天(18 日)晚間報道,在優步(Uber),AI 已經滲透到員工工作的許多環節,從回答 Slack 裡的內部問題,到替乘客行程中聯繫客服時收到的消息撰寫回復。6 名近期遭裁員的員工透露,過去幾個月,AI 在工作中的使用範圍明顯擴大,其中一些人甚至會通過提示詞讓 AI 完成相當一部分任務。

智譜 ZCode 被質疑“偷傳代碼”:官方回應稱問題已修復,將開源代碼庫、引入第三方審查
作者:清源 責編:清源 評論: 感謝網友 咩咩洋 的線索投遞!9 月 18 日消息,針對社區中有關代碼庫數據上傳的討論,智譜旗下編程產品 ZCode 今天(18 日)通過智譜官方群組向受影響用戶致歉,併發布回應稱已第一時間完成自查,相關問題目前已經修復。

月之暗面遞表之後,Kimi 的成色要被驗算三遍
舒澤品牌手記2026.09.18 18:16 · 來自浙江全文4982字00:00 / 14:05Anthropic 的 30 萬次指控,會成為招股書的第幾頁?文 | 舒澤品牌手記9月17日,月之暗面發佈了一套金融行業解決方案。按官方披露,中信建投、中金公司、易方達等數十家金融機構已經在用 Kimi 處理投研建模、風險排查和盡調材料——研究人員把管理層報表、審計報告和盡調文件交給 Kimi,拿回一份可以繼續調整假設的 Excel 模型。同一天,深圳商報記者就港股上市進展、股東架構調整等事項向月之暗面發去採訪函。

Calibre上手 AI 互動寫作:電子書管理器搖身變成"文字冒險遊戲引擎"
這個遊戲默認藏而不發,不會跟著 Calibre 啟動就冒出來。用戶得主動在"首選項 — 工具欄和菜單"裡把它請到主工具欄,才算真正激活。它的玩法很清晰:由 AI 在後臺搭起並掌管一個虛構世界,用戶通過不斷輸入文字來推著故事往前走,等於把"讀電子書"這件事,翻轉成了"和 AI 一起寫故事"。