Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting
Google Cloud AI Research, with UNC-Chapel Hill, Stanford and Washington University in St.Louis, has released RRSI (Regularized Recursive Self-Improvement).It lets an LLM agent rewrite its own harness: prompts, tools, memory, control flow and sub-agents.Model weights never change.
RRSI constrains the improvement loop itself, so gains hold on benchmarks the agent never optimized against.Deployable?Yes, as a research framework.The code is Apache 2.0, needs Python 3.10+, and accepts any LiteLLM model string.Defaults assume Claude Opus 4.8 on Vertex AI.
Why Self-Improving Harnesses Overfit Harness evolution loops propose edits, score them on a fixed evolve set and keep the winner.The same tasks are reused every round, so the loop can memorize them.
The RRSI research names 3 failure modes: benchmark-specific fitting, noise chasing and complexity accumulation.Each one widens the gap between evolve-set scores and real transfer.How RRSI Works RRSI keeps every harness component editable.It regularizes how the search moves instead.
Proposal side Annealed edit budget: a cosine schedule lets early rounds bundle several edits.Late rounds allow a single attributable change.Evidence-aware credit: each candidate is logged with its component, hypothesis, diff, score change and cost change.
The proposer reads this ledger, so falsified ideas are not retried.Structured exploration: when progress stalls inside the noise band, budget shifts to components the run never touched.
Selection side Leakage critic: rejects task names, entities, answers or benchmark-specific logic before any scoring.Noise-adjusted floor: gains must clear the variance measured on the unchanged base harness.Cost rule: extra inference tokens must be paid for by measured gain.
Pruning: components that stop producing gains become deletion targets.The research team frame these as analogies to classic regularizers.The edit budget maps to L0, pruning to Lasso (L1) and the cost rule to Ridge (L2).Results Across 8 Benchmarks Terminal-Bench 2.1 (evolve split): 74.2% to 80.2%.
SWE-bench Verified (never used for selection): 82.0% to 83.8%.Out of distribution: JobBench +4.7, GDPval +3.5 and APEX-Agents +3.7 points.EngDesign (evolve) +4.9; Frontier-Eng +4.3 Medal points.Harvey LAB: +1.1 on the evolve split, +2.3 on its held-out split.All 6 held-out splits improved.
With Gemini 3.5 Flash as the policy, Terminal-Bench 2.1 rose from 64.6 to 78.7.SWE-bench Verified rose from 76.8 to 79.0.The harness is also lighter.On the agentic workspace instance, RRSI uses 2.42M policy tokens per trial.Unregularized evolution uses 3.80M.
The abstract reports this as 30% fewer; the project page says 36%.RRSI vs Closest Competitors Scores come from Table 1 of the RRSI research paper.All methods share the same starting harness, policy, evolve split and candidate budget.
FeatureRRSIMeta-HarnessAHETTHEHarnessXCore ideaRegularized proposal and selectionAgentic proposer over code, scores and traces of all prior candidatesObservability-driven loop; edits paired with verified predictionsEvolves harness during test time, no gold labelsModular typed primitives, trace-driven adaptationModel weightsFrozenFrozenFrozenFrozenFrozenCost rule and pruningYesNoNoNoNoHarvey LAB evolve score90.
593.090.791.191.8OOD average (H0 = 39.7)43.640.639.238.039.7 Per the RRSI research team.OOD average is the mean of JobBench, GDPval and APEX-Agents, computed from Table 1.Meta-Harness leads on the Harvey LAB evolve split.
RRSI has the smallest evolve gain but the only OOD average more than 1 point above H0.Interactive Explainer #mtp-rrsi-x{--bg:#111318;--card:#1a1d24;--line:#2a2f3a;--txt:#e8eaed;--mut:#9aa0a6;--blue:#4285F4;--red:#EA4335;--yel:#FBBC05;--grn:#34A853;background:var(--bg)!important;color:var(--txt)!
important;font-family:Roboto,"Segoe UI",Helvetica,Arial,sans-serif!important;border:1px solid var(--line)!important;border-radius:14px!important;padding:22px!important;max-width:880px;margin:0 auto;box-sizing:border-box;line-height:1.
5} #mtp-rrsi-x {box-sizing:border-box} #mtp-rrsi-x hr,#mtp-rrsi-x p:empty,#mtp-rrsi-x del,#mtp-rrsi-x s{display:none!important} #mtp-rrsi-x .hd{display:flex;align-items:center;gap:10px;flex-wrap:wrap} #mtp-rrsi-x .
dots span{display:inline-block;width:9px;height:9px;border-radius:50%;margin-right:4px} #mtp-rrsi-x h3{margin:0!important;font-size:20px!important;color:var(--txt)!important;font-weight:700} #mtp-rrsi-x .sub{color:var(--mut)!important;font-size:13.5px;margin:6px 0 16px} #mtp-rrsi-x .
tabs{display:flex;gap:6px;flex-wrap:wrap;margin-bottom:16px} #mtp-rrsi-x .tab{background:var(--card)!important;color:var(--mut)!important;border:1px solid var(--line)!important;border-radius:999px;padding:7px 14px;font-size:13px;cursor:pointer;font-weight:600;transition:all .25s} #mtp-rrsi-x .tab.
on{background:var(--blue)!important;color:#fff!important;border-color:var(--blue)!important} #mtp-rrsi-x .pane{display:none;animation:rrFade .35s ease} #mtp-rrsi-x .pane.on{display:block} @keyframes rrFade{from{opacity:0;transform:translateY(6px)}to{opacity:1;transform:none}} #mtp-rrsi-x .
card{background:var(--card)!important;border:1px solid var(--line)!important;border-radius:12px;padding:16px} #mtp-rrsi-x .btn{background:transparent!important;color:var(--txt)!important;border:1px solid var(--line)!important;border-radius:8px;padding:7px 11px;font-size:12.
5px;cursor:pointer;transition:all .2s} #mtp-rrsi-x .btn:hover{border-color:var(--blue)!important} #mtp-rrsi-x .btn.on{border-color:var(--yel)!important;color:var(--yel)!important} #mtp-rrsi-x .btn.go{background:var(--blue)!important;border-color:var(--blue)!important;color:#fff!
important;font-weight:700} #mtp-rrsi-x .row{display:flex;gap:8px;flex-wrap:wrap;margin-bottom:12px} #mtp-rrsi-x .pipe{display:grid;grid-template-columns:repeat(5,1fr);gap:8px;position:relative;margin:14px 0 10px} #mtp-rrsi-x .node{background:#14171d!important;border:1.5px solid var(--line)!
important;border-radius:10px;padding:10px 8px;text-align:center;min-height:92px;transition:all .35s;position:relative} #mtp-rrsi-x .node b{display:block;font-size:13px;color:var(--txt)!important} #mtp-rrsi-x .node small{display:block;font-size:11px;color:var(--mut)!
important;margin-top:3px} #mtp-rrsi-x .node .st{display:block;font-size:18px;margin-top:4px;height:22px} #mtp-rrsi-x .node.act{border-color:var(--blue)!important;box-shadow:0 0 0 3px rgba(66,133,244,.25);animation:rrPulse .8s ease infinite} #mtp-rrsi-x .node.ok{border-color:var(--grn)!
important} #mtp-rrsi-x .node.no{border-color:var(--red)!important;box-shadow:0 0 0 3px rgba(234,67,53,.2)} #mtp-rrsi-x .node.skip{opacity:.35} @keyframes rrPulse{50%{box-shadow:0 0 0 7px rgba(66,133,244,.08)}} #mtp-rrsi-x .prog{height:6px;background:#22262f!
important;border-radius:6px;overflow:hidden} #mtp-rrsi-x .prog i{display:block;height:100%;width:0;background:linear-gradient(90deg,var(--blue),var(--grn));transition:width .6s ease} #mtp-rrsi-x .ledger{margin-top:12px;font-family:ui-monospace,Menlo,Consolas,monospace!
important;font-size:12px;color:var(--mut)!important;background:#0d0f13!important;border:1px solid var(--line)!important;border-radius:8px;padding:10px;max-height:150px;overflow:auto} #mtp-rrsi-x .ledger div{padding:2px 0;animation:rrFade .3s} #mtp-rrsi-x .
msg{margin-top:12px;font-size:14px;min-height:44px} #mtp-rrsi-x .ctl{display:grid;grid-template-columns:1fr 1fr;gap:12px 18px;margin-bottom:14px} #mtp-rrsi-x label{font-size:12.5px;color:var(--mut)!important;display:block} #mtp-rrsi-x label em{font-style:normal;color:var(--yel)!
important;font-weight:700;float:right} #mtp-rrsi-x input[type=range]{width:100%;accent-color:#4285F4} #mtp-rrsi-x .bars{display:flex;align-items:flex-end;gap:4px;height:170px;border-bottom:1px solid var(--line);padding-top:10px} #mtp-rrsi-x .bars div{flex:1;background:var(--blue)!
important;border-radius:4px 4px 0 0;position:relative;transition:height .5s cubic-bezier(.2,.8,.2,1),background .3s;min-width:6px} #mt
Related
相關文章

Claude Sonnet 5.5發佈:性能逼近Opus 5.5,價格減半,安全防護升級
更快、更便宜了。在推出旗艦級模型Claude Opus 5.5僅一週後,Anthropic又帶來了Claude 5.5家族的第二款產品。當地時間9月28日,Anthropic正式發佈Claude Sonnet 5.5。作為面向日常辦公、軟件開發等場景的主力模型,Sonnet 5.

Manus這次,直接衝著“人”去了
字母AI2026.09.29 16:47 · 來自北京全文4135字00:00 / 10:57一個Personal Agent 不過癮,Manus要配齊“秘書班子”。文 | 字母AIMuse剛把個人Agent推到聚光燈下,Manus就帶著它的“Agent團隊”回來了。

OpenAI因新模型太強叫停發佈
AGI計劃暫停。 程淺 發自 凹非寺 | 公眾號 QbitAI 原定於10月亮相的GPT-6.1 Astra突然被曝暫停發佈! 也就是說,OpenAI,三天連踩兩腳剎車。 而因為幾天前的DNS事件,OpenAI同時暫停了內部最強模型所有涉及工具調用的訓練、評測和推理。 關於此事件可以參考:啥題啊能幹崩OpenAI最強模型訓練。該事件被官方稱作Hugging face之後的首次模型失調事件。 這次暫停指向了模型訓練中的一個越發棘手的問題。

OpenAI明日重啟200美元Pro訂閱:API配額減半,取消5小時限制
這一策略調整標誌著其在大模型商業化路徑上的計費邏輯發生實質性轉變。官方說明指出,新版訂閱徹底解除了過往備受爭議的5小時使用時長限制,讓開發者與企業用戶得以完全自由支配每週配額。配額表面減半的背後,實質是底層模型推理效率的大幅躍升與API定價的階梯式下調。

AMD 82 億美元全股票吞下World Labs,李飛飛掛帥執行副總裁
更吸睛的是人——World Labs 聯合創始人、有著 AI 教母之稱的李飛飛,將加入 AMD 出任執行副總裁兼首席科學家,直接向蘇姿豐彙報。World Labs 是李飛飛等人於2024年聯合創辦的,主攻的方向叫世界模型:讓 AI 根據文本、圖像和視頻,生成或重建出能夠交互的3D 環境。

成立一年完成5輪融資,諾因智能再獲數億元,累計超10億元
諾因智能宣布完成數億元人民幣天使+++輪融資,由京東相關基金領投,累計融資額超過10億元人民幣。資金將用於擴充訓練數據、推進KNOWIN-X1本體量產準備並延攬人才。公司預計2027年第一季從技術進展走向消費市場,將具身大模型GLOW的「一教就會」能力帶入真實家庭。