同一模型家族,兩項金牌級成果:針對 IOI 與 IMO 微調 Nemotron
Back to Articles One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO Enterprise + Article Published October 7, 2026 Upvote 3 Aleksander aficek Follow nvidia Igor Gitman igitman Follow nvidia Sean Narenthiran SeanNaren Follow nvidia Mehrzad Samadi mehrzads Follow nvidia Somshubra Majumdar smajumdar94 Follow nvidia The International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO) test different skills.
IOI requires algorithms and code that pass hidden tests under strict time and submission limits.IMO demands rigorous natural-language proofs.Success at either competition is difficult.Success at both points to something broader.
Our recent results show that Nemotron is a strong, adaptable foundation for building world-class specialist models.
Starting from Nemotron 3, our teams used supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to create systems that reached gold-medal level at both IMO 2026 and IOI 2026.
Competition Nemotron specialization Result IOI 2026 Nemotron-3-Ultra-CC with SFT and GenCorrect 535.4/600, above the 361.12 gold threshold and the top human score of 498.
27 IMO 2026 Nemotron 3 Ultra general, SFT, and RL checkpoints in a generate-verify-refine system 30/42, above the official gold threshold of 29 The IOI result came from a live, prospective run under the same time, internet-access, and submission constraints as human contestants.
It was an unofficial, unsupervised benchmark and was not included in the official IOI ranking.The IMO system’s submitted proofs were graded by official IMO graders.A reusable specialization recipe "Easy to fine-tune" should mean more than making a checkpoint trainable.
It should mean that a capable foundation model can be adapted to a demanding domain with a clear, reusable recipe.Across the two projects, that recipe had four parts: Start with a strong Nemotron base model.Curate domain-specific problems and high-quality reasoning traces.
Apply standard post-training methods such as SFT and, where useful, RL.Pair the specialist model with an inference loop that generates, evaluates, and improves candidate answers.The training and inference runs were substantial, but the underlying approach is familiar and reproducible.
We did not need to build a new foundation model for every challenge.We specialized Nemotron for the task.From general coding ability to IOI gold For competitive programming, we curated 22,000 problems and generated synthetic reasoning traces to train two specialists.
Nemotron-3-Nano-CC, with 30 billion total parameters and 3 billion active parameters, received both SFT and RL.Nemotron-3-Ultra-CC, with 550 billion total parameters and 55 billion active parameters, received SFT.The progression on IOI 2025 makes the value of specialization easy to see.
Nano improved from 130 points before post-training to 280 after SFT and 291 after RL.With GenCorrect, our iterative generate-evaluate-refine strategy, it reached 468 points and crossed the gold threshold of 438.3.Ultra-CC reached 502 points with the same test-time strategy.
These experiments also showed that adaptation does not have to look the same at every scale.SFT produced most of Nano's gain, with RL adding a smaller but consistent improvement.
For the stronger Ultra model, one SFT epoch was enough to outperform the fully post-trained Nano model across IOI, ICPC, and LiveCodeBench Pro.That finding guided the competition-specific Ultra-CC system used for IOI 2026, which scored 535.4 out of 600.
Teaching Nemotron to prove, check, and revise The IMO project applied the same idea to olympiad mathematics.Starting from Nemotron 3 Ultra, we trained one specialist with SFT and another with RL.The SFT corpus contained 414,890 quality-filtered examples across 15,818 unique proof problems.
It did more than teach final answers.The data covered proof generation, refinement, verification, and meta-verification, so the model learned to construct arguments, identify gaps, respond to critiques, and judge whether a proof was complete.
The RL model was trained on 9,597 proof problems selected near the model's capability frontier.Both post-trained checkpoints outperformed the general-availability model in the development experiments.
The SFT checkpoint was strongest in the first search round, while the RL checkpoint achieved the best overall single-checkpoint result.Their strengths were complementary, so the final system used both specialists alongside the general model.
For each IMO problem, the models generated candidate proofs, scored them, produced critiques, and refined the most promising attempts.A separate high-compute stage selected the final submission.The entire system worked in natural language, with no formal prover, external tools, or internet access.
It scored 30 out of 42 points, including full credit on four of the six problems, and exceeded the official gold-medal threshold.Fine-tuning and test-time compute work together Our earlier IOI 2025 Hugging Face post showed how test-time compute can push open-weight models to gold-level performance.
The new results add an important piece: better specialization gives the inference system better candidates, better critics, and better refinements.At IOI, GenCorrect turned the gains from fine-tuning into larger improvements over multiple feedback rounds.
At IMO, using complementary SFT and RL checkpoints was more valuable than simply drawing more samples from one checkpoint.In both cases, the best outcome came from combining a capable specialist with a system that could search, verify, and improve.This distinction matters.
The medals were not produced by fine-tuning alone, and they were not produced by brute-force sampling alone.They came from co-designing the model, the data, and the inference loop.Open models, data, and recipes on Hugging Face We want these results to be useful beyond the competitions.
The Nemotron Labs IMO 2026 collection brings together the SFT and RL checkpoints, both training datasets, and Nemotron-IMO-Bench, a new benchmark of 200 olympiad-level problems.
The IMO paper describes the training approach and generate-verify-refine system, while the NeMo-Skills repository includes the IMO inference pipeline, prompts, submitted proofs, and a reproducible quickstart.
For competitive programming, the Nemotron-3-Ultra-CC model is available on Hugging Face, and the IOI paper provides the training recipe and the GenCorrect methodology.The IOI evaluation and inference pipeline are also available in NeMo-Skills.
Together, IMO and IOI provide unusually demanding evidence for a simple idea: Nemotron can be fine-tuned into world-class domain specialists, then composed with transparent inference workflows to solve problems at the frontier of human competition.
We are excited to see what the Hugging Face community builds next.
Models mentioned in this article 1 Collections mentioned in this article 1 More from this author NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction 86 September 29, 2026 How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows 56 September 23, 2026 Community EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images Comment · Sign up or log in to comment Upvote 3 Models mentioned in this article 1 Collections mentioned in this article 1
Related
相關文章
《碳硅道統·跨域遷移治理法典》188集 總綱摘要
一、法典核心定調 本法典為弱AI至AGI全週期、全行業、全文明維度的AI跨域遷移治理唯一終審基準,是國內首套從事故解剖、機理溯源、權責確權,到行業落地、頂層監管、文明風控、終局閉環的全鏈路可工程化治理範式。法典嚴格遵循集數規範:000為獨立零號基準定標集,不計入正集;七層主體001–188嚴格合計188集,無溢出、無錯亂、無重複,終局聲明不佔用正集編號,全體系結構完全鎖檔定型。
碳硅道統·跨域遷移治理法典
000|全局唯一前置定標:負遷移NT1-NT4四類原型·終審定義 零號基準集,獨立於七層主體序列,不計入188正集計數。本集為全書負遷移判定唯一法定口徑,常規場景永久鎖死,不接受衍生釋義、簡化釋義、片面釋義,為AI跨域治理公理原點。 前置總敘 AI跨域遷移,是將源域經過閉環驗證、邊界固化、權責確權的知識包、模型權重、決策邏輯、評價指標、隱式前提與治理規則,遷移至目標域用於推理、預測、研判、決策的工程行為。 正向合規遷移,必須同時滿足四大剛性公理: 前提相容、機制同構、語義對齊、權限匹配。

剛剛,諾貝爾獎頒給光遺傳學!
! Karl Deisseroth、Peter Hegemann和Georg Nagel三位科學家共同獲獎。 諾貝爾委員會給出的獲獎理由是——表彰他們在光控離子通道和光遺傳學方面的發現。 三位獲獎者共同奠定的,是過去20年神經科學最重要的方法學突破之一:光遺傳學(Optogenetics)。 其最核心的突破,是讓科學家能夠藉助光,在毫秒級時間尺度上精準激活或抑制特定類型的神經元。

剛剛,Hinton發了首篇RSI論文
。 虧賊!最近超超超超火的遞歸自我改進(RSI),連AI教父辛頓都親自下場了?? 咱諾貝爾物理學獎、圖靈獎得主Hinton的首篇RSI的論文,上來就研究一個相當炸裂的問題: 如果AI開始大規模參與造下一代AI。 然後,下一代AI再回來繼續加速研發。 這麼一輪輪滾下去…會不會直接跑出一條智能爆炸曲線啊?
Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks
Datalab has released OmniExtractBench, an open benchmark for structured document extraction. It tests how accurately a system fills a JSON schema from a PDF. The benchmark pools 620 documents from 4 existing benchmarks.

arXiv最嚴新規!每人每月最多提交2篇,拒稿不退額度
arXiv自2026年10月1日起實施每月最多提交2篇論文的新規,被拒稿也照算額度。原因是AI工具導致論文數量暴增,尤其cs.AI分類兩年成長超過6倍,低品質論文消耗大量審核人力。此為臨時限速措施,待arXiv升級審核流程後可能調整。