Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize
An agent harness is the code around a model: execution loop, tools, context, state, recovery, and verification.Per the Terminal-Bench 2.1 leaderboard, GPT-5 solves 35.2% of tasks inside Terminus 2 but 49.6% inside Codex CLI with identical weights.Most benchmarks keep that harness fixed.
HarnessDev proposed by team of researchers from ByteDance Seed, Singapore University of Technology and Design, Georgia Institute of Technology, M-A-P, and TokenWave.AI, flips the target: the artifact under evaluation is the runnable harness the model writes, not the answer it produces.
2 stages: Creation and Evolution In Creation, every creator receives the same weak seed: passive file, search, and process primitives plus result and trajectory writers, with no loop, planner, verifier, retry, or stopping rule.Unmodified, it scores 0 everywhere.
The creator gets a task-family spec, a short design tutorial, and 1 to 3 development cases, builds a full harness, and the harness is frozen before hidden tasks.
In Evolution, the creator starts from its own frozen Creation code harness and revises it using execution feedback from a fixed set of 100 SWE-bench Pro tasks and all 89 Terminal-Bench 2.1 tasks.
Each official candidate must complete both evaluations as a pair, with a budget of 10 pairs and at most 2 five-task probes between pairs.Every official version is later scored on 630 held-out SWE-Pro instances the creator never sees.
Harnesses are graded on capability (task success) and efficiency (executor tokens, with creator tokens excluded).Setup 6 creator LLMs were tested: Opus 4.8, GPT-5.5, Gemini 3.1 Pro, DeepSeek V4 Pro, Qwen 3.7 Max, and Seed 2.0 Pro, working inside Claude Code 2.1.177 (GPT-5.5 used Codex 0.144.3).
Creation spans 4 domains and 5 benchmarks totaling 2,207 instances: SWE-bench Pro public split (731), Terminal-Bench 2.1 (89), MLE-bench (75), EQ-Bench3 (46), and BrowseComp (1,266).Each creator builds 3 harnesses per benchmark, reported as avg@3.
Self-Eval runs each harness with its creator; Unified-Eval runs all with Gemini 3.1 Pro.Creation results Under Self-Eval, Opus 4.8 posts the highest average score at 67.8 against a human-engineered reference of 86.2.The gap depends on domain: Code: Opus 4.8 reaches 69.3 on SWE-Pro versus the 80.
0 reference.Gemini 3.1 Pro leads Terminal-Bench at 68.8 versus 88.8.Search: the widest gap.The best BrowseComp score is 52.6 (GPT-5.5) against a 92.2 reference.Writing: Opus 4.8 scores 84.6 on EQ-Bench3, above the 83.7 reference.ML experimentation: Opus 4.8 (32.9) and Gemini (32.4) beat the 24.
0 MLE-bench reference.The SWE-Pro, Terminal-Bench, and BrowseComp references are external results from OpenAI’s GPT-5.6 report, not re-runs.Code volume did not predict quality: the 18 code harnesses added 17,111 net lines, yet Gemini added the fewest (1,006) and led Terminal-Bench.
Self-test count barely correlated with score (Spearman 0.13 to 0.26); revision calls reached 0.57.Much generated machinery is inert.Of 108 code component instances, 72 trigger in real runs and 18 never fire, all of them state and memory.
11 of 18 harnesses define a State class, yet no checkpoint event appears across 26,679 trajectories.124 of 587 writing features are dead code.Cost and executor transfer MLE-bench token use varied roughly 19-fold.GPT-5.5 hit a 19.1 medal rate with 29.3M tokens while DeepSeek V4 hit 19.6 with 208.4M.
Swapping the executor to Gemini reshuffled rankings: Qwen gained 17.6 points on BrowseComp and 12.9 on MLE-bench, while Opus 4.8’s SWE-Pro score fell from 69.3 to 33.0, partly because one harness hard-coded a 120-step limit around its original executor.
The Opus search harness’s duplicate-query rate jumped from 10.1% to 88.2% after the switch.Evolution results 9 lineages (5 self-runtime, 4 fixed-Gemini) produced 73 official versions and 64 adjacent switches.All 5 self-runtime creators improved on held-out tasks, from +1.43 to +4.44 points (mean +3.
11).Under fixed Gemini, only Opus improved; GPT-5.5 regressed 10.32 points.Progress was not monotonic.Of 64 switches, 8 regressed on both benchmarks, 16 on one, 27 gained only within the noise band, and 2 showed clear positive evidence.A single commit can vary by about ±4.75 pair-score points.
Feedback and held-out scores moved in the same direction only 34 of 64 times (53.1%), and only 2 of 9 declared final versions were held-out optimal.Of 169 new functions or classes, 25 have no caller.The clearest win: Opus 4.
8 noticed 99 of 100 runs reported success while only 48 passed, traced it to premature completion, and added a completion gate.Failure diagnosis was otherwise the weakest step: the dedicated trajectory interface was called only twice.Interactive explainer (function(){window.
addEventListener("message",function(e){var d=e.data;if(d&&d.type==="mtp-harnessdev-resize"&&d.height){var f=document.getElementById("mtp-harnessdev-frame");if(f)f.style.height=d.height+"px";}});})(); Key Takeaways HarnessDev scores the harness a model builds, not the answer it returns.
Self-built harnesses match or beat references on writing and ML experimentation but trail badly on code and search.Harness quality is executor-specific; Opus 4.8 drops from 69.3 to 33.0 on SWE-Pro under Gemini.
Evolution gains are small, noisy, and only 34 of 64 changes point the same way on held-out tasks.Much generated state and memory code never executes.Check out the Paper and Project Page.
Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?
Connect with us The post Can LLMs Engineer Their Own Agent Harness?ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize appeared first on MarkTechPost.
Related
相關文章

無問芯穹與華環電子簽署戰略合作,共同探索國產異構算力AI基礎設施新方向
無問芯穹與華環電子簽署戰略合作協議,雙方將結合各自在AI軟體平台、網路通信與硬體研發的優勢,共同探索國產異構算力基礎設施的協同方案。此次合作聚焦於智算中心解決方案及「Token工廠」新模式,目標是推動計算、網路與AI原生基礎設施深度融合,為AI規模化應用提供高效穩定的支撐。
Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs
Jina AI, part of Elastic, has released jina-ocr-v1, an end-to-end visual document parser. It takes PDFs, scans, tables, charts or invoices and returns clean Markdown in 1 pass. The model has 3.

優步全球範圍裁員 10%,被裁員工稱 AI 已大舉滲透日常工作
作者:清源 責編:清源 評論: 9 月 18 日消息,據《商業內幕》今天(18 日)晚間報道,在優步(Uber),AI 已經滲透到員工工作的許多環節,從回答 Slack 裡的內部問題,到替乘客行程中聯繫客服時收到的消息撰寫回復。6 名近期遭裁員的員工透露,過去幾個月,AI 在工作中的使用範圍明顯擴大,其中一些人甚至會通過提示詞讓 AI 完成相當一部分任務。

智譜 ZCode 被質疑“偷傳代碼”:官方回應稱問題已修復,將開源代碼庫、引入第三方審查
作者:清源 責編:清源 評論: 感謝網友 咩咩洋 的線索投遞!9 月 18 日消息,針對社區中有關代碼庫數據上傳的討論,智譜旗下編程產品 ZCode 今天(18 日)通過智譜官方群組向受影響用戶致歉,併發布回應稱已第一時間完成自查,相關問題目前已經修復。

月之暗面遞表之後,Kimi 的成色要被驗算三遍
舒澤品牌手記2026.09.18 18:16 · 來自浙江全文4982字00:00 / 14:05Anthropic 的 30 萬次指控,會成為招股書的第幾頁?文 | 舒澤品牌手記9月17日,月之暗面發佈了一套金融行業解決方案。按官方披露,中信建投、中金公司、易方達等數十家金融機構已經在用 Kimi 處理投研建模、風險排查和盡調材料——研究人員把管理層報表、審計報告和盡調文件交給 Kimi,拿回一份可以繼續調整假設的 Excel 模型。同一天,深圳商報記者就港股上市進展、股東架構調整等事項向月之暗面發去採訪函。

Calibre上手 AI 互動寫作:電子書管理器搖身變成"文字冒險遊戲引擎"
這個遊戲默認藏而不發,不會跟著 Calibre 啟動就冒出來。用戶得主動在"首選項 — 工具欄和菜單"裡把它請到主工具欄,才算真正激活。它的玩法很清晰:由 AI 在後臺搭起並掌管一個虛構世界,用戶通過不斷輸入文字來推著故事往前走,等於把"讀電子書"這件事,翻轉成了"和 AI 一起寫故事"。