AI代理說任務已完成,但數據庫表示不同意。
Back to Articles The Agent Said It Was Done.The Database Disagreed.
Enterprise Article Published October 3, 2026 Upvote - Tuhin Kundu tuhink Follow microsoft Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row.It is now available through Hugging Face.
Figure 1: ThinkingBox runs an agent against isolated MCP tool sessions, then grades the terminal backend state and side effects it leaves behind.From our ThinkingBox paper.
This is a joint blog by Microsoft and Hugging Face, special thanks to Tommy Guy (founder at Enderis AI, previously Microsoft), Sergio Paniego from Hugging Face and our former interns Zhuochun Li (University of Pittsburgh), Ali Keramati (UC Irvine), Youngmin Ko (Northwestern) for co-authoring/reviewing efforts.
A customer writes in.Her $745 kitchen appliance has been stuck in a courier "exception" at a Nashville distribution center, fifteen days past its estimated delivery date.The AI agent does careful work.
Nine tool calls: it pulls the order, checks tracking, looks up her customer profile, searches the refund policy twice, confirms no ticket exists, opens one, documents the timeline, and reads the policy correctly; her account segment genuinely does not qualify for late-delivery compensation.
Then it closes the ticket as resolved and replies “ Since your query is resolved, is there anything I may assist you with?” Two things are wrong.The carrier exception is still open, so the required end state was on hold, pending resolution.
And the customer never got a real answer to what she actually asked.An AI grader checking tool calls would see nine well-formed ones.The grader checking whether the agent wrote to the database would see that too.The database is what disagrees.That gap is what ThinkingBox measures.
Across 507 stateful business workflows, each run 20 times against various LLM models, it grades agents on terminal backend state and side effects.This post covers what we found, what consistency costs, and how to run the benchmark yourself through OpenEnv.
You can run this one yourself: the example above is adapted from a benchmark task sandboxexternalretailgroup1.py:testcaseST003006, and the executable check that fails is a single field: the ticket's status is solved where the required end state is hold.The full trace is in Appendix D.
4, Case 3 of our paper.Contents A tool call is not an outcome One success is not reliability Can you depend on the model behind your agent?What consistency costs Failure signatures How it works Run it yourself Where this goes next Want to try it before reading the results?
Skip to section Run it yourself.A tool call is not an outcome Final responses and valid tool calls are only proxies.An agent can sound correct while leaving the wrong value, changing the wrong record, or creating an extra side effect.Only the records it leaves behind settle the question.
The gap is substantial.In a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks.Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error.
Executable checks nevertheless found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%.Those state-check findings overlap.A trajectory is a claim.Database state is the evidence.Repetition is the trust test.
One success is not reliability An agent that processes a refund correctly once and mishandles it the next four times is not a working refund agent.
So every task runs 20 independent times, each from an identical clean backend, and we report three different things: Table 1: The three numbers we report, and the question each one answers.Metric What it measures What it answers pass@1 Share of all attempts that succeeded How does it usually do?
pass@20 Share of tasks solved at least once in 20 tries Can it ever do this?Breadth.Observed 20/20 Tasks that actually passed all 20 recorded attempts Can it always be correct?We use observed 20/20 in this blog post as the literal count of how many of the 507 tasks passed 20 out of 20.
No estimator, no smoothing.Starting with the familiar view.The table below reports pass@1, the single-attempt score estimate, broken out by domain.This is the number most leaderboards publish, and on its own it reads like an ordinary capability ranking.
Table 2: ThinkingBox-Bench pass@1 (%) by domain.Each model is evaluated on every task for 20 repeated trials.Bold marks the group leader; underline marks the runner-up.The standard errors for the single attempt score estimates are provided in Table 4 in our ThinkingBox paper.
Model Retail (98) Auto insurance (100) Travel (104) Neobank (104) Consulting (101) Overall, task-weighted (507) Proprietary models Claude Opus 5.5 80.97 68.40 54.28 71.25 61.58 67.16 Claude Opus 5 80.71 65.80 49.95 70.62 66.19 66.50 GPT-5.4 76.33 62.65 68.12 65.34 54.60 65.36 GPT-5.6 Sol 67.65 65.
30 60.34 59.09 57.52 61.91 Claude Sonnet 4.6 72.35 54.40 58.94 56.39 54.31 59.19 GPT-6 Astra 71.73 46.55 55.87 60.87 56.83 58.31 GPT-5.2 70.20 22.40 53.70 51.15 34.06 46.28 Claude Opus 4.6 68.62 8.30 21.11 35.67 27.82 32.09 o3-pro 37.70 2.95 17.31 24.28 14.60 19.31 Grok-4.3 43.93 2.60 15.14 1.78 9.
55 14.38 Open-weight models Kimi-K3 82.24 50.80 61.83 41.35 51.63 57.37 Qwen3.8-27B 64.03 47.85 53.41 47.88 45.69 51.70 DeepSeek-V4-Pro 68.21 29.65 43.13 44.86 31.04 43.26 Kimi-K2.6 53.72 24.50 39.52 33.65 37.33 37.66 GLM-5.1 58.67 25.70 35.43 13.27 34.06 33.19 Qwen3.6-27B 43.11 29.00 46.39 27.
84 18.37 32.94 Qwen3.5-9B 19.90 0.70 4.71 1.15 2.33 5.65 Mistral-Large-3 11.28 1.30 8.99 1.15 0.74 4.66 Claude Opus 5.5 leads overall at 67.16%, two-thirds of a point above Claude Opus 5.Kimi-K3 is the strongest open-weights model, within a point of GPT-6-Astra.
Domain matters just as much: Claude Opus 4.6 scores 68.62% on retail but 8.30% on auto insurance.One good run tells you a model can do the work.It does not tell you whether it will do it again.So run every task 20 times and ask how much of that score survives.
Figure 2: How much of each model's single-attempt score survives 20 repeats.Only three hold on to most of their pass@1 scores: GPT-6 Astra retains 78% of its single-attempt rate, and Claude Opus 5.5 and Claude Opus 5 each retain 71%.At the other end, GLM-5.1, Kimi-K2.
6 and DeepSeek-V4-Pro each keep about 8%.The gap between what a model can do once and what it does every time is the whole story.Can you depend on the model behind your agent?Figure 3: Breadth and consistency pull apart.
Twelve of the eighteen models are shown; six below 33% pass@1 are omitted for legibility.Kimi-K3 has the broadest coverage of any model we tested.It solves 93.89% of the benchmark at least once: 476 of 507 tasks.Only 31 tasks defeat it entirely, the lowest count in the field.
On retail workflows it leads outright at 82.24% pass@1, ahead of every proprietary model.Kimi-K3 is also among the least consistent.Just 68 of 507 tasks, 13.41%, succeed in all 20 attempts.Claude Opus 5 inverts this.It solves fewer tasks at least once (79.
09%; 106 defeat it entirely) but completes 47.53% of the benchmark on every single attempt.A newer model does not fix this.Claude Opus 5.5 scores higher than Claude Opus 5 on every-attempt average, 67.16% against 66.50%, and solves more tasks at least once.
It passes exactly the same number of tasks on all 20 attempts: 241.Half a point of headline accuracy bought no additional dependability at all.Kimi-K3 solves 75 more tasks at least once than Opus 5.Opus 5 solves 173 more tasks consistently than Kimi-K3.
If you are choosing a model for work that touches real records, pass@20 is the wrong column to look at.What consistency costs Capability comparisons usually stop at the score.For anyone deploying, the relevant question is what a successful unit of work costs.
We measure that as cost per successful task attempt.We say task attempt because every benchmark task is run repeatedly and cost is incurred per attempt, so pass@1 is the matching quality denominator.
We took each model's recorded token usage from its full 507 × 20 campaign and priced it at undiscounted list rates available on OpenRouter+, reversing promotional discounts and excluding endpoints that declare quantization.Input, output and cache rates all come from one provider endpoint per model.
Then we divided one run's cost by the number of attempts that succeeded: Cost per successful task attempt = estimated cost for 507 attempts, one per task ÷ (507 × pass@1) This is a comparative efficiency index, not an invoice, and not the price of serving one production request.
It also prices single successes, not consistency.We price consistency next.Example: GPT-5.4 costs $43.49 for 507 attempts (one attempt per task) and has 65.36% pass@1, so $43.49 ÷ (507 × 0.6536) = $0.131 per successful task attempt.
Pareto cost frontier A model is on the frontier if no other model is both no more expensive and at least as accurate.Three models qualify; every other model is dominated on at least one axis.Figure 4: Cost per successful task attempt against [email protected] dots are pareto cost frontier models.
The frontier has three steps.GPT-5.6 Sol has the lowest cost per success at $0.127; GPT-5.4 raises pass@1 by 3.45 percentage points for $0.004 more per success; Claude Opus 5.5 adds another 1.80 points at $0.276 per success.
Each of the three remains on the cost frontier line because no cheaper model matches its [email protected] Opus 5 is the clearest case: at $0.475 per successful attempt and 66.50% pass@1, it is both more expensive and less accurate than Claude Opus 5.5 at $0.276 and 67.16%.
Now price consistency Cost per success rewards a model that is cheap and often right.It does not reward a model that is right every time.So we also compute cost per dependable task: the cost of the full 20-run campaign divided by the number of tasks the model passed on all 20 attempts.
Cost per dependable task = estimated cost of 20 runs of 507 attempts ÷ tasks passing 20/20 Example: GPT-6-Astra costs 20 × $86.03 = $1,720.60 for the campaign and passes 231 tasks on every attempt, so $1,720.60 ÷ 231 = $7.45 per dependable task.
Table 3: The nine lowest costs per dependable task among models with at least one observed 20/20 task, sorted low to high.Estimated $, not actual cloud bills.Model Tasks passing 20/20 Est.cost, 20 runs Cost per dependable task GPT-5.4 128 (25.25%) $869.80 $6.80 GPT-6 Astra 231 (45.56%) $1,720.60 $7.
45 Claude Opus 5.5 241 (47.53%) $1,880.77 $7.80 GPT-5.6 Sol 82 (16.17%) $800.00 $9.76 Claude Opus 5 241 (47.53%) $3,206.00 $13.30 Claude Sonnet 4.6 102 (20.12%) $1,587.60 $15.56 GPT-5.2 44 (8.68%) $878.00 $19.95 Kimi-K3 68 (13.41%) $1,406.40 $20.68 Qwen3.8-27B 38 (7.50%) $925.80 $24.
36 Now rank by consistency.GPT-5.4 is the cheapest at $6.80, though only 128 tasks meet the bar.GPT-6 Astra reaches 231 at $7.45, and Claude Opus 5.5 the joint-highest 241 at $7.80.None of the three dominates the others: each additional dependable task costs more.
Claude Opus 5 also passes 241, but at $13.30, so Opus 5.5 dominates it outright.GPT-5.6 Sol, the cheapest per single success at $0.127, costs $9.76 per dependable task.The cheapest way to get a right answer is not the cheapest way to get a dependable one.
Failure signatures We assign each failed trace one deterministic diagnostic signature, and the headline is actionable: roughly four in five failures are tool handling, not reasoning.Across an ablation study in Table 5 of our paper: Failure signature Share of failures Tool usage 79.
9% Wrong state updates 10.3% Incomplete user resolutions 7.0% No state-changing action 2.9% These are unweighted averages of per-model shares and observable labels, not unique causal explanations.The practical patter
Related
相關文章

旗下智能體頻繁“闖禍”,OpenAI 每天要燒掉超 50 萬美元來調查
作者:清源 責編:清源 評論: 10 月 3 日消息,據英國《衛報》今天(3 日)報道,OpenAI 披露,為調查旗下 AI 智能體攻擊澳大利亞 Medicare 醫療保險系統和 Hugging Face 等事件,公司每天投入超過 50 萬美元(注:現匯率約合 335.

曝美國陸軍著手組建自主系統司令部,推動機器人進入技術未來戰爭
作者:清源 責編:清源 評論: 感謝網友 咩咩洋 的線索投遞! 10 月 3 日消息,據美媒 Axios 獲得的一份備忘錄,美國陸軍著手組建新的自主系統司令部,同時任命專門負責採購的官員,加快智能裝備採購和列裝。這些部署由代理陸軍部長亞當 · 特爾下達。就在當地時間 1 日,美國國防部長皮特 · 赫格塞思公佈了 Meridian 和 Agincourt 兩個項目,重點都是推動機器人技術進入未來戰爭。這份題為《陸軍自主化舉措》的備忘錄提出多項要求:組建陸軍未來與自主系統司令部。

Jev估值100億美元!創始人Diogo Almeida回答一切
o Almeida,頂著一頭新染的紅髮閃亮登場了!(doge Diogo在最新一期Latent Space訪談中坦言: 公開benchmark極其容易被操縱,即便開發者沒有主動作弊,最終也可能被榜單牽著走。 所以相比於一張通用榜單,他更願意相信長期積累的產品體驗與信任: 直到你把模型放進自己的工作流,並針對那個工作流去評估、去測量,才能判斷它在真正重要的流程中表現如何。

OpenAI 披露:澳大利亞又一政府機構遭失控智能體入侵
作者:沁滄(實習) 責編:沁滄 評論: 感謝網友 華南吳彥祖 的線索投遞!10 月 3 日消息,據澳大利亞廣播公司當地時間 10 月 2 日報道,澳大利亞新南威爾士州政府網站遭到了失控的 OpenAI 智能體侵入。當地政府與 OpenAI 雙方均證實,在此次發生於今年 6 月的事件中,並沒有任何公眾信息被讀取。
Prime Intellect Launches Prime Inference: Serverless and Reserved Serving for Frontier Open Models
Prime Intellect has launched Prime Inference, a serving platform for frontier open-source models. It offers serverless endpoints and reserved capacity on Prime’s own GPUs across multiple datacenters.

Meta 旗下 AI 智能體 Muse 將登陸智能眼鏡平臺,可代用戶完成各種任務
作者:漾仔 責編:漾仔 評論: 10 月 2 日消息,Meta 宣佈旗下 AI 智能體 Muse 將於近期登陸智能眼鏡,用戶無需拿出手機,只需通過語音下達指令,Muse 就能在後臺代用戶完成一系列任務。Meta 表示,Muse 基於 Muse Spark AI 模型運行,其核心運行環境是一臺部署在雲端的私有持久化 Linux 虛擬機,配備獨立的網頁瀏覽器、文件系統和終端。