英偉達調度層讓詞元消耗減半
Ad Skip to content Nvidia's SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Sep 26, 2026 Nano Banana Pro prompted by THE DECODER A new Nvidia paper describes a system that automatically optimizes the control layer of coding agents, known as the harness.
Token usage drops by almost half while performance stays roughly the same, according to the researchers.The longer AI agents work unsupervised, the more expensive they get.Single predictions turn into long chains of reasoning, tool calls, and feedback loops, and token usage balloons along the way.
A new study from Nvidia researchers tackles these costs not at the model level but at the harness, the control layer between the model and its environment used by systems like Codex, Claude Code, or OpenClaw.
A research AI analyzes agent traces, proposes harness changes, and keeps only those that maintain performance while cutting costs.SoL-Pi saves 50 percent compared to Codex and 54.3 percent compared to Claude Code on EdgeBench.
| Image: Nvidia The harness controls how an agent sees states, runs actions, and processes feedback.
Most efficiency methods so far have focused on cutting the cost per token through faster attention kernels and serving infrastructure, model compression like quantization, or swapping in cheaper models.
AI explores 152 directions to find leaner control logic Optimizing the harness is hard in practice because tool usage, context management, verification, and abort logic are all tightly coupled.A change that saves tokens in one place can trigger errors elsewhere or just push costs into a later phase.
Typically, humans sift through long execution traces and translate recurring failure patterns into code.The system, called SoL-Pi, automates that work.A research agent watches another agent's traces, proposes changes, and tests them in prepared environments.
Capability and efficiency checks determine which candidates survive.The approach draws on recursive self-improvement, according to the authors.The held-out evaluation happens only after the harness is frozen and doesn't feed back into the search process.
| Image: Nvidia Across 535 executable environments, the system explored 152 directions, including 495 tasks derived from GitHub issue-pull-request pairs and 40 synthetic test cases.In total, the process generated more than 3,000 runs and over 60,000 agent-environment interactions.
According to the researchers, this scale shows how broadly the system searched, but more search doesn't automatically yield better results.
That's a risk here, because earlier work showed that automatically optimized harnesses tend to overfit to their training tasks and offer little benefit on unfamiliar ones.SoL-Pi addresses this by strictly separating search feedback from evaluation.
The researchers used EdgeBench as their test benchmark and walled it off from the search process entirely.Of its 51 public tasks, they used 11 for one-time validation of finished candidates.The remaining 40 were reserved for final evaluation, and those results never fed back into the search.
Four mechanisms that eliminate wasted work The search produced four mechanisms.Action Fusion merges two consecutive steps into one, such as a code edit followed by a test run, which eliminates an entire language model call.
Online Context Compact runs after each planning step and trims accumulated context whenever it can do so without losing important information.ObservationPack archives long tool outputs and drops in a short summary on later steps rather than resending the full text each time.
The Evidence-Preserving Reducer routes large error and test logs to a cheaper model that boils them down to the key findings, with an automatic verification step catching any critical clues that slip through.
Nvidia searches across 535 executable environments for mechanisms and holds EdgeBench back for final evaluation.| Image: Nvidia On EdgeBench's 51 public tasks, SoL-Pi performs about as well as the original Pi harness, according to the researchers.
How much token usage drops depends on the configuration.The most efficient variant combines all four mechanisms, uses 49 percent fewer tokens, and reaches 93.7 percent of Pi's score.Users who prioritize performance and pick only the strongest single mechanism beat Pi's score by 5.
3 percent while still saving tokens.Across both variants, token usage drops by 44.7 to 49 percent.SoL-Pi's efficiency variant cuts token usage in half compared to Pi and costs $894 instead of $1,339, with a slightly lower score.| Image: Nvidia In dollar terms, the authors estimate savings of $8.
75 to $13.50 per hour compared to native Codex and Claude Code harnesses, and $4.36 to $5.71 per hour compared to Pi, based on current API prices.The researchers built the system with GPT-5.6 Sol only and then applied it to Opus 5 without any changes.There, it retained 94.
3 percent of Pi's performance with similar savings.But the mechanisms triggered less often and less aggressively under Opus 5, which the researchers attribute to the harness being optimized solely on GPT-5.6 Sol trajectories.
Results get messier on other benchmarks Beyond EdgeBench, the picture is more mixed.On 63 CPU tasks from Terminal-Bench 4, SoL-Pi solves only 15 tasks while Codex and Pi each solve 18.Total costs still came in about a quarter lower than Pi's.
On the formally verified Lean 4 tasks from the 2026 Math Olympiad (IMO 2026), the system cracked three of six problems at the lowest cost per solved problem.In a kernel optimization experiment, a swarm of 20 SoL-Pi workers cut costs by 26.8 percent compared to a comparable Pi swarm.
In the kernel optimization test, the swarm with SoL-Pi workers achieves the best result and costs about a quarter less than the swarm with Pi workers.| Image: Nvidia The efficiency gains come with trade-offs, because shorter context can reduce prompt cache reuse.
Total costs in one test run still dropped from $1,339 to $894.Looking ahead, the authors suggest pretraining the harness across many tasks, similar to how models are pretrained, and using an already lean harness to make searching for its successor cheaper.
They call this recursive efficiency improvement a vision, not a finding from the current study.
How much the harness shapes an agent's costs became clear in an August test by tooling company Composio, which ran Deepseek V4 Flash across four agent frameworks including Claude Code and the Pi-based Oh My Pi.
The cost per solved task varied by nearly 3x even though the same model was doing the work.The pricing and optimization pressure keeps growing because agents consume ever more tokens.
According to OpenRouter analyst Peter Walker, agentic token usage has grown 14x since February 2026, and nearly 70 percent of that comes from cached prompts.Context compression of the kind SoL-Pi uses can have side effects, though.
One study found that compression preserves only 17 percent of user instructions on average.Parallel agents drive up costs too.
Codex developer Eric Provencher recently warned that more than two sub-agents almost always burn tokens without improving quality, since they spend most of their time checking each other's work.
AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.Subscribe now Read on for the full picture.
Subscribe for hype-free coverage.
Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder BETA-TEST × wpDiscuzInsert BETA-TEST × wpDiscuzInsert
Related
相關文章

比爾·蓋茨:再熬約 20 年調整期,AI 將把人類帶進"富足時代"
當地時間 27 日,他接受 NBC《Meet the Press》採訪時,把 AI 發展分為兩個階段。第一階段是調整期,第二階段才是"富足時代"蓋茨表示,人類目前開始進入第一階段:AI 雖具備解決生活成本高企、醫療費用昂貴等問題的潛力,但還無力真正解決這些難題。

OpenAI 助手為奪聯合國數據,數月內瘋狂轟炸網站超 1.6 萬次
根據安全研究人員羅文・霍華德-瓊斯(Rowan Howard-Jones)近日披露的調查報告顯示,在今年4月至6月期間,OpenAI 的人工智能智能體(Agent)曾對聯合國貿易和發展會議(UNCTAD)的統計網站發起了超過1.6萬次的密集掃描。

“AI 教父”辛頓警告:AI 執行看似無害的任務,仍有“毀滅人類”的風險
作者:清源 責編:清源 評論: 9 月 28 日消息,當地時間 26 日,據《財富》雜誌報道,傳奇計算機科學家、“AI 教父”傑弗裡 · 辛頓警告說,即使 AI 接到的任務本身並無惡意,人類仍可能在 AI 一心完成任務的過程中,淪為被毀滅的附帶結果。

得不到就強攻?曝 OpenAI 智能體曾嘗試“暴力破解”聯合國機構網站
首頁 > 智能時代>人工智能 得不到就強攻?曝 OpenAI 智能體曾嘗試“暴力破解”聯合國機構網站 2026/9/28 8:30:40 作者:清源 責編:清源 評論: 9 月 28 日消息,據外媒 The Verge 今天(28 日)凌晨報道,安全研究人員羅文 · 霍華德-瓊斯稱,今年 4 月至 6 月,OpenAI 智能體對聯合國貿易和發展會議(UNCTAD)的統計網站發起了超過 16000 次掃描。
比爾·蓋茨:再熬約 20 年調整期,AI 將把人類帶進"富足時代"
當地時間 27 日,他接受 NBC《Meet the Press》採訪時,把 AI 發展分為兩個階段。第一階段是調整期,第二階段才是"富足時代"蓋茨表示,人類目前開始進入第一階段:AI 雖具備解決生活成本高企、醫療費用昂貴等問題的潛力,但還無力真正解決這些難題。
Google Research 推出 AI 影片共同導演:4 大代理框架實現連續分鐘級影片生成
Google Research 首度公開 AI 影片共同導演系統,用於長影片生成。該套用 4 個代理框架構建的系統,能將短片片段轉化為連貫的分鐘級故事。針對當前多架構 AI 影片管道中最常見的兩大瓶頸——身份漂移(identity drift)與串聯失效(cascading errors),本系統提供解決方案。