何夕2077生成式AI

需求驅動智能體訓練任務合成框架問世

2026年7月23日 00:00

重點摘要

Computer Science > Software Engineering arXiv:2607. 14186 (cs) [Submitted on 15 Jul 2026 (v1), last revised 22 Jul 2026 (this version, v4)] Title:NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synth。

站內 AI 整理稿

Computer Science > Software Engineering arXiv:2607. 14186 (cs) [Submitted on 15 Jul 2026 (v1), last revised 22 Jul 2026 (this version, v4)] Title:NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs Authors:Jiarong Zhao, Zhikai Lei, Zhiheng Xi, Rui Zheng, Hang Yan, Jie Zhou, Qin Chen, Liang He View a PDF of the paper titled NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs, by Jiarong Zhao and 7 other authors View PDF Abstract:Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual substrate engineering, each new domain demands a bespoke pipeline, and the resulting task distributions often reflect substrate biases rather than real-world demand. We introduce NexForge, a requirement-driven framework that takes high-level capability requirements as input and synthesizes diverse, executable agent tasks and expert trajectories for SFT. NexForge first investigates real-world demand to construct representative scenarios and task profiles, then performs distribution-aware compilation to generate task directives. For each directive, NexForge automatically retrieves or constructs the required files, dependencies, and runtime configurations, and finally synthesizes expert rollouts and produces training trajectories. Without domain-specific infrastructure, NexForge produces 3. 6K terminal and 2K office tasks, improving Qwen3. 5-35B-A3B Base from 22. 0\% on Terminal-Bench 2.

Related

相關文章

AI大模型進入「無限戰爭」

2026年7月17日凌晨,馬斯克在X平台簡短回應了一則關於Kimi K3的評測,只用了一個詞:「Impressive」。這款由月之暗面推出的新模型,參數規模達到2.8萬億,是當時全球最大的開源權重模型,在Artificial Analysis的智能指數中排名第三,僅次於Claude Fable 5和GPT-5.6 Sol。隔天,馬斯克又補了一句,說xAI正在訓練的2萬億參數模型「可能超過Kimi」。月之暗面則在微博上幽默回應:「歡迎加入『2萬億+』俱樂部」。

剛剛

GPT-5.6推翻近30年數學猜想,全程對話公開:提示詞只有58個單詞???

GPT-5.6 Pro 最近在數學界掀起波瀾,繼先前推翻雅可比猜想後,這回又盯上了圖論領域一個懸而未決近三十年的難題——Dinitz-Garg-Goemans 猜想。一位名叫 Dmitry Rybin 的研究者,在與 AI 的互動過程中,只用了四條提示詞,總共 58 個英文單詞,就讓模型端出了一個足以推翻這道猜想的反例。整個過程沒有複雜的提示詞工程,沒有密密麻麻的數學公式,通篇幾乎就是「繼續研究」「繼續找」「給我一個完整反例」這類直白的催促。

剛剛
量子位生成式AI

WAIC最狠展臺打爆工業「深水區」!它石智航首發具身原生大腦AWE 3.5,具身Scaling全面釋放

它石智航在WAIC發布具身原生基座模型AWE 3.5,能讓機器人無需切換參數即可執行多種工業任務,並在展會現場展示1:1還原的線束自動化產線。該模型從預訓練階段就整合視覺、語言與動作,成為首個完整打通預訓練到後訓練範式的原生模型,大幅降低新任務所需數據量,宣告具身Scaling全面釋放。

剛剛

K3之後,Kimi為什麼急著上市?

月之暗面在發布2.8萬億參數的K3開源模型後,五天內接連宣布暫停新用戶註冊並啟動港股IPO計畫,顯示其急於在技術與商業化雙重拐點鎖定資本市場。Kimi的3億美元年度經常性收入(ARR)中API收入佔比逾七成,證明中國大模型靠API賺錢的模式可行,但300人團隊與算力瓶頸仍是上市後需面對的挑戰。

剛剛