智能體編程超人類
Computer Science > Artificial Intelligence arXiv:2604.
02721 (cs) [Submitted on 3 Apr 2026 (v1), last revised 5 Aug 2026 (this version, v3)] Title:GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning Authors:Ornith Team: Xiaoya Li, Guoyin Wang, Songqiao Su, Chris Shum, Jiwei Li View a PDF of the paper titled GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning, by Ornith Team: Xiaoya Li and 4 other authors View PDF HTML (experimental) Abstract:Competitive programming remains one of the last few human strongholds in coding against AI.
The best AI system to date still underperforms the best humans competitive programming: the most recent best result, Google's Gemini~3 Deep Think, attained 8th place even not being evaluated under live competition conditions.
In this work, we introduce GrandCode, a multi-agent RL system designed for competitive programming.
The capability of GrandCode is attributed to two key factors: (1) It orchestrates a variety of agentic modules (hypothesis proposal, solver, test generator, summarization, etc) and jointly improves them through post-training and online test-time RL; (2) We introduce Agentic GRPO specifically designed for multi-stage agent rollouts with delayed rewards and the severe off-policy drift that is prevalent in agentic RL.
GrandCode is the first AI system that consistently beats all human participants in live contests of competitive programming: in the most recent three Codeforces live competitions, i.e.
, Round~1087 (Mar 21, 2026), Round~1088 (Mar 28, 2026), and Round~1089 (Mar 29, 2026), GrandCode placed first in all of them, beating all human participants, including legendary grandmasters.
GrandCode shows that AI systems have reached a point where they surpass the strongest human programmers on the most competitive coding tasks.Comments: Tech Report; Pre-print Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2604.02721 [cs.AI] (or arXiv:2604.02721v3 [cs.
AI] for this version) https://doi.org/10.48550/arXiv.2604.
02721 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Jiwei Li [view email] [v1] Fri, 3 Apr 2026 04:26:56 UTC (9,653 KB) [v2] Mon, 13 Jul 2026 00:56:35 UTC (9,653 KB) [v3] Wed, 5 Aug 2026 18:24:09 UTC (9,653 KB) Full-text links: Access Paper: View a PDF of the paper titled GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning, by Ornith Team: Xiaoya Li and 4 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.
AI < prev | next > new | recent | 2026-04 Change to browse by: cs References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...
Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.
ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?
) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.
AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?
) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)
Related
相關文章

阿里雲棲大會亮出最激進AI藍圖: 5 萬億至 10 萬億參數模型在途,自研真武V900 芯片算力翻三倍
首席執行官吳泳銘宣佈,公司正準備訓練參數規模介於5萬億至10萬億之間的新一代模型,這是阿里迄今最激進的AI發展計劃,也意味著它要從雲計算服務商進一步蛻變為全棧人工智能平臺。按照披露,這款仍在研發中的新模型將成為通義千問(Qwen)系列的重要後續產品。

消息稱高瓴創投合夥人嚴文韜加入DeepSeek,擔任CFO
此次任命意味著DeepSeek成立三年來長期空缺的CFO職位正式補齊,也顯示其公司治理和財務管理體系進一步完善。資料顯示,嚴文韜1991年出生,畢業於復旦大學,2013年至2020年間先後任職於騰訊投資、H Capital,2020年加入高瓴創投。

Qwen4已投入訓練,阿里公佈10萬億參數模型演進路線
Qwen-Image-2.1、HappyShrimp 1.1 及HappyOyster 2.0-Preview也同步推進。開源生態方面,一個月內Qwen3.8 相關模型下載量已超 5600 萬次,衍生模型超 1900 個;截至目前,阿里已開源 460 多個千問模型,整體下載量超 30 億次,衍生模型超 30 萬個。

工業AI正在走向產線,但規模化仍有卡點丨ToB產業觀察
Leo張ToB雜談2026.09.22 13:51 · 來自江蘇全文3920字00:00 / 11:45工業軟件的玩法,正在從“賣一套標準產品”轉向“幫客戶把AI能力長在自己身上”。過去,一家造船廠給排一份為期三個月的生產計劃,要把時間和空間約束都考慮進去,需要兩個員工,花費超過半個月的時間。

阿里研究員透露Qwen4.5後模型將擴展至5-10T參數
阿里巴巴在雲棲大會上宣布新一代架構的Qwen4模型已開始訓練,未來Qwen4.5、Qwen5等版本參數規模將擴展至5至10兆。大會同時展示大模型遞歸自我改進技術已應用於訓練與推理,並推出多款影像、音樂、語音及全模態模型新版本。阿里也宣布下代視頻生成模型預計11月發布,並將語音模型落地於手機、AI眼鏡等終端裝置。

阿里公佈全模態模型新進展,Qwen4和下代視頻模型均在訓練中
阿里巴巴在2026雲棲大會上公布多項大模型進展,包括基於新架構的Qwen4已開始訓練,未來參數規模將擴展至5到10萬億。多模態方面,影片生成模型Wan3.0在評測中取得雙榜第一,下一代影片模型預計11月發布;語音、影像、音樂及世界模型等也同步升級。此外,Qwen3.8-Max透過自我進化技術,在零人工參與下持續迭代,整體下載量已超過30億次。