英偉達極速小模型正式亮相
Ad Skip to content Nvidia's open-weight Nemotron 3.5 Lightning prioritizes speed over maximum intelligence Matthias Bastian View the LinkedIn Profile of Matthias Bastian Aug 11, 2026 Nano Banana Pro prompted by THE DECODER Nvidia's new Nemotron 3.
5 Lightning is a compact open-weights model that matches OpenAI's gpt-oss-120b on intelligence benchmarks with a quarter of the parameters while delivering the fastest inference speeds in its class.Nvidia has released Nemotron 3.5 Lightning, the first model in its new Nemotron 3.5 lineup.
The model directly succeeds the Nemotron 3 Nano 30B A3B and keeps its hybrid Mamba-Transformer architecture, with 31.6 billion total parameters and only 3.6 billion active at any given time.
According to the independent benchmarking platform Artificial Analysis, the model scores 24 on the Intelligence Index, a nine-point jump from its predecessor (15).
That puts Lightning on par with OpenAI's gpt-oss-120b (24) and just behind Nvidia's own Nemotron 3 Super (26), which is about four times larger.The smartest small models in the same size class, like Qwen3.6 35B A3B (32) and Meta's new Muse Glimmer (35), still hold a clear lead.Ad Nemotron 3.
5 Lightning scores 24 on the Artificial Analysis Intelligence Index, tying with gpt-oss-120b.| Image: Artificial Analysis Nvidia is targeting a different spot on the efficiency frontier with Lightning.
In pre-release tests using the final NVFP4 weights, the model hits nearly 670 tokens per second, the highest measured throughput among all compared models and almost twice as fast as Google's Gemini 3.5 Flash-Lite (386 tokens/s).A task from the Intelligence Index takes about 0.
5 minutes to complete, while Qwen3.6 35B A3B needs around 3.5 minutes, and Gemma 4 31B takes roughly 5.8 minutes.AdDECDIncontent-1 At 669 tokens per second, Lightning is the fastest model in the comparison.
Despite generating a similar number of tokens per task as its predecessor Nemotron 3 Nano, it delivers much better results.| Image: Artificial Analysis Proprietary models still dominate the overall efficiency frontier.Gemini 3.
5 Flash-Lite scores 37 on the Intelligence Index with a similar time per task, and GPT-5.6 Luna (max) reaches 52 points in under two minutes.Agentic benchmarks show the biggest gains The biggest improvements show up in agentic benchmarks, according to Artificial Analysis.
On GDPval-AA v2, Lightning reaches an Elo rating of 824, a 334-point gain over Nemotron 3 Nano.That beats both gpt-oss-120b (800) and the larger Nemotron 3 Super (698).On Terminal-Bench v2.1, the score jumps from 7 to 24.3 percent, nearly matching gpt-oss-120b at 26.2 percent.
Ad On agentic benchmarks, Lightning surpasses both gpt-oss-120b and the larger Nemotron 3 Super with an Elo rating of 824.| Image: Artificial Analysis Nvidia ships the model under the permissive OpenMDW-1.1 license, positioning it as a high-throughput workhorse for agent-based pipelines.
Artificial Analysis reports that Nvidia worked with partners like CodeRabbit and Harvey on post-training to boost performance in specific domains.Availability Nvidia provides the model in both BF16 and NVFP4 weights.
The NVFP4 variant also scores 24 on the Intelligence Index with minimal quality loss compared to the higher-precision version, according to Artificial Analysis.The reasoning model handles text only and supports a context window of one million tokens.
Weights are available now, and serverless inference is offered by DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe, among others.AdDECDIncontent-2 Nvidia's push for efficiency over size isn't new.
In a widely discussed paper last year, its researchers argued that models under 10 billion parameters can handle most agent workloads as well as 70- to 175-billion-parameter models at one-tenth to one-thirtieth the cost.Nemotron 3.5 Lightning has 31.6 billion parameters but activates only 3.
6 billion per step, putting it in the same lightweight class.At nearly 670 tokens per second, it also beats gpt-oss-120b and the larger Nemotron 3 Super on agentic benchmarks, making it the clearest product-level proof of that thesis yet.
Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now Source: Nvidia | Artificial Analysis | Hugging Face BETA-TEST × wpDiscuzInsert BETA-TEST × wpDiscuzInsert
Related
相關文章

仇太深,奧特曼炮轟A社“反人類”
量子位·2026年08月24日 16:01“Codex名字沒起好,我也沒用明白” 《奧特曼天降正義,小嘴抹毒暗諷A社“反人類”》!!這個標題有沒有港媒內味兒了(doge),您先別笑,這還真是奧特曼在最新採訪裡的大致意思。雖然沒有點名,但話裡話外就差直接報Dario Amodei的身份證號了。

消息稱字節整合 AI 生產力:TRAE、釦子併入豆包,將推統一辦公品牌“豆包工作”
作者:沁滄(實習) 責編:沁滄 評論: 感謝網友 HH_KK 的線索投遞!8 月 24 日消息,據智能湧現消息,字節跳動對旗下的辦公 AI 產品完成了一輪團隊整合:TRAE、釦子(Coze)團隊將整體併入豆包體系,其中 TRAE Work、釦子將與豆包在工作場景的產品能力進行整合;TRAE IDE 及 CLI 將作為豆包品牌下的編程產品線持續發展。

AI 大模型周榜:國產 glm-5.3-max 首秀闖入綜合榜前 15,kimi-k3-max 衝進前十
本週AI大模型Arena排行榜出現新面孔,智譜AI的glm-5.3-max首次入榜即拿下綜合榜第13名。月之暗面的kimi-k3-max排名持續攀升,成功擠進綜合榜前十,位列第10名。兩款國產大模型雙雙寫下佳績,成為本週關注焦點。

DeepSeek Harness來了:AI開始製造AI了?
DeepSeek Harness 正式推出,這項新工具被視為 AI 發展的重要里程碑,可能讓 AI 系統具備自主開發或優化其他 AI 的能力。外界關注此技術是否象徵 AI 開始「製造」AI,並可能加速人工智慧的進化與應用。目前相關細節與實際影響仍待進一步觀察。
神秘“牛來”大模型上線即登頂 背後廠商至今未揭曉
近日,一款代號為Ox Alpha的匿名AI模型在OpenRouter悄然上線,短時間內調用量迅速攀升,成功衝至平臺榜首,並刷新了該平臺的單日模型用量紀錄。不過,儘管表現驚豔,Ox Alpha背後的開發主體至今仍未揭曉。
具身智能資本熱浪再起,小鵬機器人首輪估值突破63億美元
本輪融資由IDG資本領投,高榕創投參投,並獲騰訊和阿里巴巴作為戰略投資者共同參與,小鵬集團仍保持控股地位。四家投資方均將該輪視為其在具身智能領域迄今披露的最大規模單筆投資之一。IDG資本指出,小鵬人形機器人代表當前國內產業領先水平,已具備與海外頭部企業全球競爭的技術實力。