Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation

2026年9月11日 06:01
站內 AI 整理稿

Training an LLM to call tools reliably requires datasets that pair user queries with correct tool-use chains.Producing that data at scale has been slow and expensive.A team of researchers from Google, the University of Tokyo, RIKEN AIP, and Tohoku University introduce ToolGrad.

The research work inverts the usual pipeline: build a verified tool chain first, then write the query.Gemma-3 models fine-tuned on 500 samples of the resulting data reach scores that sit alongside frontier proprietary models on the Berkeley Function Calling Leaderboard.Is it deployable?Yes.

The code is Apache-2.0, the ToolGrad-500 dataset and the 1B, 4B, and 12B models are on Hugging Face, and there is a PyPI package.The problem with query-first generation Prior pipelines such as ToolBench and ToolACE follow a query-first recipe.

The system samples a pool of APIs, asks an LLM to invent a plausible user instruction, and then dispatches a depth-first search (DFS) agent to find a tool-use path that satisfies it.The search has no guarantee of success.

When it dead-ends, the compute spent on exploration is wasted, and the sample is discarded.The paper frames this as distilling valuable trajectories from a complex and often failing agent exploration, which is inherently inefficient.ToolGrad reverses the order.

It first constructs a ground-truth tool-use chain by actually executing APIs, then annotates that chain with a matching user query.An explicit, working chain is far less ambiguous than a hypothetical prompt, so the chain-to-query step takes a single LLM call.

Four modules in a loop Each iteration runs four modules in sequence: API Proposer narrows a sampled set of APIs down to a few candidates that could extend the current workflow.API Executors run those candidates in parallel and produce detailed execution reports.

API Selector reviews the reports, picks the single best-performing call, and appends it to the workflow.Its directional feedback is the textual gradient.LLM Updater rewrites the synthetic user query and AI response so they match the new API set.

Repeating the loop yields one sample: a user query, a verified API workflow, and the final response.The repository’s default configuration runs 10 iterations over 50 sampled APIs per workflow.

Generation efficiency on ToolBench The research team evaluated data generation on the ToolBench API database, which contains 16,000+ real-world APIs, and compared ToolGrad against ToolBench’s DFS-based query-first approach.According to the research paper: Pass rate rose from 63.8% (DFS) to 99.

8% (ToolGrad).Ground-truth tool uses per sample rose from 2.1 to 3.4, meaning longer chains.Tool-use steps per sample fell from 34.3 to 20.0.LLM invocations per sample fell slightly, from 64.5 to 63.9.The 0.

2% failure case occurred when the agent could not get a successful response from 3 selected APIs across all 10 iterations and saved an empty sample.Interactive explainer (function(){var f=document.getElementById("mtp-toolgrad-frame");window.addEventListener("message",function(e){if(e.data&&e.data.

type==="mtp-toolgrad-resize"&&e.source===f.contentWindow){f.style.height=e.data.height+"px";}});})(); BFCL results with Gemma-3 The researchers generated ToolGrad-500, a 500-sample dataset built with Gemini 2.5 Flash-Lite, and used it to post-train Gemma-3 at 1B, 4B, and 12B parameters.

They evaluated on the Berkeley Function Calling Leaderboard, which uses a tool set that differs from ToolBench, making it an out-of-distribution test with unseen tools.Findings reported by the authors: Fine-tuning on ToolGrad-500 improved tool-use scores at every parameter size.

ToolGrad-12B scored 83.1, compared with Gemini 2.5 Pro at 83.2, Claude 4.5 Opus at 82.8, and GPT-5 at 74.4, as measured at the time of publication.The 12B student outperformed Gemini 2.5 Flash-Lite, the teacher model that generated its training data.

ToolGrad-12B led open tool-use specialists including ToolACE and Hammer-2.1-7B.The repository’s reproduction scripts target BFCL V1 and V2 through a customized fork, run inference in a vLLM Docker image, and were verified on a single NVIDIA A100 40GB.

Key Takeaways ToolGrad flips tool-use data generation: verify the chain first, write the query second.Pass rate jumps from 63.8% to 99.8% on ToolBench, with longer chains and fewer tool steps.Only 500 samples lift Gemma-3-12B to 83.1 on BFCL, next to Gemini 2.5 Pro at 83.2.

The student model beats its Gemini 2.5 Flash-Lite teacher.Check out the Paper, GitHub Page, and Google Research Blog.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?

now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Google Research Releases ToolGrad: Answer-First Framework Hits 99.

8% Pass Rate for Tool-Use Data Generation appeared first on MarkTechPost.

Related

相關文章

量子位生成式AI

無問芯穹與華環電子簽署戰略合作,共同探索國產異構算力AI基礎設施新方向

無問芯穹與華環電子簽署戰略合作協議,雙方將結合各自在AI軟體平台、網路通信與硬體研發的優勢,共同探索國產異構算力基礎設施的協同方案。此次合作聚焦於智算中心解決方案及「Token工廠」新模式,目標是推動計算、網路與AI原生基礎設施深度融合,為AI規模化應用提供高效穩定的支撐。

4 分鐘前
IT之家生成式AI

優步全球範圍裁員 10%,被裁員工稱 AI 已大舉滲透日常工作

作者:清源 責編:清源 評論: 9 月 18 日消息,據《商業內幕》今天(18 日)晚間報道,在優步(Uber),AI 已經滲透到員工工作的許多環節,從回答 Slack 裡的內部問題,到替乘客行程中聯繫客服時收到的消息撰寫回復。6 名近期遭裁員的員工透露,過去幾個月,AI 在工作中的使用範圍明顯擴大,其中一些人甚至會通過提示詞讓 AI 完成相當一部分任務。

3 小時前
鈦媒體生成式AI

月之暗面遞表之後,Kimi 的成色要被驗算三遍

舒澤品牌手記2026.09.18 18:16 · 來自浙江全文4982字00:00 / 14:05Anthropic 的 30 萬次指控,會成為招股書的第幾頁?文 | 舒澤品牌手記9月17日,月之暗面發佈了一套金融行業解決方案。按官方披露,中信建投、中金公司、易方達等數十家金融機構已經在用 Kimi 處理投研建模、風險排查和盡調材料——研究人員把管理層報表、審計報告和盡調文件交給 Kimi,拿回一份可以繼續調整假設的 Excel 模型。同一天,深圳商報記者就港股上市進展、股東架構調整等事項向月之暗面發去採訪函。

5 小時前

Calibre上手 AI 互動寫作:電子書管理器搖身變成"文字冒險遊戲引擎"

這個遊戲默認藏而不發,不會跟著 Calibre 啟動就冒出來。用戶得主動在"首選項 — 工具欄和菜單"裡把它請到主工具欄,才算真正激活。它的玩法很清晰:由 AI 在後臺搭起並掌管一個虛構世界,用戶通過不斷輸入文字來推著故事往前走,等於把"讀電子書"這件事,翻轉成了"和 AI 一起寫故事"。

7 小時前