NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router

2026年8月12日 06:59
站內 AI 整理稿

NVIDIA introduced open technologies for building always-on AI agents from systems of specialized models.Two artifacts shipped together.Nemotron 3.

5 Lightning is a lightweight, customizable open model built for high-volume agentic tasks, and NeMo Switchyard is an open source routing library that directs each step of an agent workflow to the most capable and efficient model available.

The problem both address is structural: long-running agents spend most of their time on tool calls, result validation, and subagent delegation, and sending every one of those steps to a frontier reasoning model adds cost and latency.

Lightning is a 30B mixture-of-experts model with 3B active parameters, built on a hybrid Mamba-2 + MoE + Attention architecture with a 1M-token context window.NVIDIA reports up to 4x faster output speed than similar-sized models, and 30% faster completion of 10,000 PinchBench tasks than Qwen3.

6 35B at comparable accuracy.Many industry players like CrowdStrike, Harvey, CodeRabbit, Fastino Labs, and Lila Sciences are already customizing it for cybersecurity, legal, coding, finance, and healthcare workloads.Is it deployable?Yes.Nemotron 3.

5 Lightning is generally available under the permissive OpenMDW-1.1 license, with open weights, training data, and recipes.NVIDIA states the model is ready for commercial use.Which companies: Anyone with a single modern GPU.NVIDIA lists single-GPU deployment on 1x DGX Spark (GB10) or 1x H100.

That puts solo developers and seed-stage startups on the same footing as enterprises.Mid-market teams can serve it from Baseten, Together AI, or Nebius; regulated enterprises can keep it fully on-premises.

Industries: Cybersecurity, legal services, software engineering, financial services, healthcare, and life sciences all appear in NVIDIA’s named customer set.

Applications: Tool calling, result validation, subagent delegation, code review routing, log triage, contract parsing, and long-context retrieval across a 1M-token window.The execution layer, not the planning layer Long-running agents spend most of their time on high-volume execution.

Tool calls, result validation, and subagent delegation dominate the token budget.Routing every one of those steps to a frontier reasoning model adds cost and latency.Nemotron 3.5 Lightning targets that execution layer.

It is a 30B mixture-of-experts model with 3B active parameters, built on a hybrid Mamba-2 + MoE + Attention architecture.Context length reaches 1M tokens.Pre-training covered more than 20 trillion tokens using an NVFP4 recipe.The model is the smallest member of the Nemotron 3 family.

Frontier models such as Nemotron 3 Ultra handle orchestration and planning, while Lightning handles the routine calls beneath them.

Where the speed comes from Two mechanisms: First, Speculative Decoding: Multi-token prediction was baked in during a dedicated pre-training stage, then improved with an MTP-boosting phase.

NVIDIA also ships two external draft models: DSpark, a semi-autoregressive drafter recommended for DGX Spark and low-concurrency data center workloads, and DFlash, which uses a lightweight block-diffusion model.Second, Quantization: An NVFP4 checkpoint ships alongside BF16.

The same checkpoint serves Blackwell and Hopper natively, and extends to Ampere through W4A16 kernels.NVIDIA reports up to 4x output speed versus similar-sized models.On PinchBench, it reports 86% accuracy while completing 10,000 tasks 30% faster than Qwen3.6 35B at comparable accuracy.

Published model card results (BF16 / NVFP4): MMLU Pro 81.94 / 81.62, GPQA Diamond 75.44 / 75.57, SWE-bench Verified 51.56 / 52.80, Terminal-Bench 2.1 24.58 / 23.46, AA-LCR 52.00 / 49.19.Recommended sampling is temperature 1.0 and top_p 0.95.

NeMo Switchyard NeMo Switchyard is an open source library that routes each step of an agent workflow to the most capable and efficient model available.

It offers tuning-free routers, including an LLM classifier with session affinity, a stage router that reads recent tool activity, and an escalation router that starts cheap and promotes on sustained difficulty.

A tunable prefill router learns from the model’s residual stream to predict which candidate will succeed.The reference server accepts OpenAI, Anthropic, and Responses API requests.Two published results: LangChain benchmarked 145 multi-turn agentic tasks.Routing between Lightning and Claude Opus 4.

8 with the escalation router cut cost 74% versus a frontier-only baseline, sending 7% of calls to the frontier model, at a roughly 6-point accuracy tradeoff.Cognition implemented staged routing in Devin Desktop.On FrontierCode Main, routing between Opus 5 and Kimi K2.7 reached 50.6% at a $3.

11 mean cost, within 2.8 points of Opus 5 accuracy at approximately 28% lower mean cost.Interactive explainer (function(){ window.addEventListener('message', function(e){ var d = e.data; if(!d || d.mtpFrame !== 'nemotron-lightning') return; var f = document.

getElementById('mtp-nemotron-lightning'); if(f && d.height) f.style.height = d.height + 'px'; }, false); })(); Key Takeaways 30B open MoE with 3B active parameters, 1M context, OpenMDW-1.1 license, commercial use permitted.

Up to 4x output speed; PinchBench 10,000 tasks completed 30% faster than Qwen3.6 35B.Speed comes from multi-token prediction plus DSpark and DFlash drafters, and an NVFP4 checkpoint.Runs on 1x DGX Spark or 1x H100, and locally via Ollama, LM Studio, llama.cpp, and Unsloth.

NeMo Switchyard cut cost 74% in LangChain’s 145-task benchmark at a ~6-point accuracy tradeoff.Try it on build.nvidia.com or OpenRouter, and download weights from Hugging Face or ModelScope.

Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?

Connect with us The post NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router appeared first on MarkTechPost.

Related

相關文章

IT之家AI Agent

消息稱 NVIDIA 考慮向 Perplexity AI 投資“數十億美元”

作者:溯波(實習) 責編:溯波 評論: 8 月 24 日消息,外媒 The Information 稍早前報道稱,NVIDIA(英偉達)考慮在 Perplexity AI 的最新融資輪中向這家人工智能初創企業投資“數十億美元”。雙方正就該交易展開磋商,還可能達成技術授權協議。

剛剛
AIBaseAI Agent

Guidelight評估五大AI實驗室,OpenAI遏制能力排名第一

該評估覆蓋Anthropic、Google、OpenAI、Meta和xAI,僅依據公開信息考察其是否建立監控、異常行為處置、第三方審計及失控模型關閉等機制。在滿分5分的評估中,OpenAI以3分排名最高,Anthropic和Meta得分最低。

5 小時前7400
何夕2077AI Agent

τ_0-VLA長程操控

τ₀-VLA是一種層次化機器人基礎模型,透過世界模型引導的測試時計算來改善長程操控任務。高層策略在決策不確定時會額外分配計算資源,搜尋替代子任務並預測其視覺結果,而低層策略則在40,115小時的異質真實世界數據上訓練。在包含13到25個步驟的真實長程任務中,封閉迴路成功率從27.5%提升至45.0%。

7 小時前
何夕2077AI Agent

智能體技能合集

VoltAgent 的智能體技能庫近期持續擴充,目前收錄的技能規模已突破千項,相關倉庫在開發社群中獲得約 31.3k 的星標關注。這個持續成長的技能庫,正逐步成為開發者快速組合與部署智能體工作流程的重要資源。 該技能庫涵蓋多種類型的 CLI 工作流程,這些流程具備高度可複用性,讓開發者不必從零開始建構每個環節,而是能直接引用既有技能來加速專案進度。隨著技能項目不斷增加,VoltAgent 生態系的應用範圍也隨之擴大,從自動化任務到複雜的指令處理,都能找到對應的現成模組。

7 小時前