NVIDIA 研究人員推出 Physis-Lang:自我進化的物理語言,在物理基準測試上讓 Cosmos 3 超越 Veo 3.1

2026年9月30日 07:32
站內 AI 整理稿

Video world models can render convincing clips that still break physics.Butter spreads like paint.Balls pass through walls.A team from NVIDIA, MIT and the University of Oxford argues the fix can come from language itself, not from extra visual, latent or numerical signals.

Their framework, Physis-Lang, treats physical language as a shared, optimizable representation.The same text drives data curation, model training and inference.On the public Physics-IQ Verified leaderboard snapshot dated September 29, 2026, Physis-Lang on Cosmos3-Super ranks first at 48.2 ± 1.4.

The Cosmos3-Nano version ranks second at 43.3 ± 1.5.A video world model can make a convincing clip and still get the physics wrong.Our researchers just released Physis-Lang, an open self-evolving framework that adds physics reasoning to video captions.

The captions explain why and how a scene unfolds.We use them to fine-tune… pic.twitter.com/agIHuIB6N2— NVIDIA AI (@NVIDIAAI) September 29, 2026 What Problem Does Physis-Lang Solve?Conventional captions describe what happens, not why.

‘Butter melts as the temperature rises’ says nothing about heat transfer or gravity.Physis-Lang adds a physicsreasoning field to each base caption.It spells out entities, causes, interactions, governing principles, temporal evolution and effects.

The pipeline also writes a scene-specific physicsnegative_prompt.This text describes likely implausible outcomes, such as a stone floating on water.It acts as negative conditioning at inference time.How Does the Self-Evolving Caption Loop Work?

The loop keeps the captioner frozen and evolves only its instruction.A GPT-5.5 captioner writes captions for a fixed 20-video development set with 273 human-verified assertions.Gemini-3.1-Pro acts as a physics-aware critic.

An evolution agent reads the scores and claim-level failures, then rewrites the prompt.The critic scores 2 dimensions: Precision: the caption is split into atomic claims, and each claim is checked against the video.

Recall: each human-curated physical assertion must be explicitly stated or entailed by the caption.Every revised prompt is validated on PhysCapBench, a new benchmark of 246 videos and 3,794 human-verified assertions.Caption F1 rose from 78.64 at iteration 1 to 87.82 at iteration 9.

The path was not smooth.Iteration 2 made captions overly cautious and dropped F1 to 76.28.Iteration 9 required every visible causal step and raised frame sampling from 2 fps to 4 fps.How Does Language-Guided Data Curation Work?A GPT-5.

5 diagnosis agent maps generated-video failures to physics categories like rigid-body motion, collision and fluid dynamics.That deficiency profile is matched against physics tags on a large video gallery.Retrieval targets physical content, not visual appearance.

The final training set holds 183K videos: 71K filtered from WISA-80K plus 112K retrieved clips.Retrieval alone added 3.01 points on average across 3 benchmarks.On VideoPhy-2, chemical and thermal processes each gained 8.00 points.(function(){var f=document.getElementById("mtp-physis-frame");if(!

f||f.dataset.mtpBound)return;f.dataset.mtpBound="1";window.addEventListener("message",function(ev){if(!ev.data||ev.data.type!=="mtp-physis-h")return;if(ev.source!==f.contentWindow)return;var h=parseInt(ev.data.h,10);if(h>200&&h<4000)f.style.

setProperty("height",h+"px","important");});})(); How Does Physis-Lang Compare With Veo 3.1?Fine-tuning uses LoRA on attention projections, with no architecture or objective change.Physis-Lang on Cosmos3-Nano versus Google's Veo 3.1: PhyGenBench: 71.04 vs 65.63 Physics-IQ Verified: 43.41 vs 34.

99 PhyGround: 69.90 vs 69.24 VideoPhy-2: 68.02 vs 68.87 on the full set, 62.36 vs 58.43 on the Hard split Gains hold across backbones: +7.05 on Wan2.1-14B, +3.24 on Cosmos3-Edge-4B, +6.22 on Cosmos3-Nano-16B and +5.02 on Cosmos3-Super-64B.

General quality held steady on VBench-I2V, where Cosmos3-Nano moved from 88.32 to 88.69.Prompting alone also helps.Physics reasoning plus negative prompts lifted a frozen Cosmos3-Nano on PhyGenBench from 61.67 to 67.29.Can It Run Without Commercial APIs?

The research team distilled the GPT pipeline into 2 Qwen3-VL-4B-Instruct models: PhysThinker-C for captioning and PhysThinker-U for prompt upsampling.On Wan2.1-14B, the commercial pipeline gave +7.05 at about $24.12K in API cost.Swapping in PhysThinker-C kept +6.76 at about $0.12K.

A fully local setup cost $0 and still added +4.76.Physis-Lang vs Closest Competitors FeaturePhysis-LangPhiZeroPhyGDPOSelf-RefinementDeveloperNVIDIA, MIT, OxfordCASIA (NLPR)Meta (ECCV 2026)Liu et al.

Core ideaSelf-evolving natural-language physics captions and negative promptsLearned discrete "physical language", reason-then-renderGroupwise DPO with VLM physics rewardsMultimodal chain-of-thought prompt refinement from VLM feedbackWhere physics entersData curation, training captions and inference promptsQwen3-VL-4B reasoner feeding a diffusion decoderPreference training on PhyVidGen-135KInference prompts onlyTraining neededLoRA SFT (prompt-only mode also helps)YesYes (DPO)No, training-freePhysics-IQ Verified43.

4140.91n/r27.20PhyGenBench71.04n/r48.9649.17VideoPhy-2 (All / Hard)68.02 / 62.36n/r59.56 / 44.9447.88 / 28.09PhyGround69.9057.85n/r58.

22Code / weights publicPaper only"Coming soon""Released soon"Paper Scores are from the Physis-Lang paper, Tables 1 to 4, run under one protocol per benchmark (PhyGenBench and VideoPhy-2 use a GPT-5.5 evaluator).Physis-Lang numbers use the Cosmos3-Nano backbone.n/r = not reported in that comparison.

Release status checked September 30, 2026.Key Takeaways Physis-Lang evolves physics captions with a critic-guided agent while the captioner stays frozen.PhysCapBench scores captions on 3,794 human-verified cause, law and effect assertions.Cosmos3-Nano with Physis-Lang beats Veo 3.

1 on 3 of 4 benchmarks.Physics prompts alone lift a frozen model by 5.62 points on PhyGenBench.No code or weights are public yet; only the paper is released.Check out the Paper, Project Page and GitHub Repo.All credit goes to the researcher of this project.

Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?

Connect with us The post NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks appeared first on MarkTechPost.

Related

相關文章

具身智能不只在地上跑:給無人機裝上機械臂,它們要去天上「擰螺栓」了 | IROS 2026

本文作者: 張豈萍 2026-09-30 14:29 導語:具身智能的研究邊界,正從地面移動操作向空中協同作業延伸。具身智能的研究邊界,正從地面移動操作向空中協同作業延伸。作者丨張豈萍 編輯丨幸麗娟 機械臂抓取、移動機器人導航和多機器人協作,是具身智能研究的常見切口,但大多仍延續“地面機器人”範式:身體默認是地面移動機器人,環境是結構化可通行空間與離散物體,任務是感知、移動、抓取、搬運等。

剛剛

美議員提出重磅法案:政府出臺安全防護機制前,AI 不得進行遞歸式自我改進

作者:清源 責編:清源 評論: 感謝網友 咩咩洋 的線索投遞!9 月 30 日消息,當地時間 28 日,據美國 CNBC 報道,代表硅谷選區的民主黨聯邦眾議員羅 · 康納準備提出一項 AI 監管法案,引入嚴格責任標準,並禁止 AI 在缺乏政府安全防護機制的情況下進行遞歸式自我改進。

1 小時前

機器人“拜師”蘇繡!最細的活,最硬的考題

機器人前瞻(公眾號:robot_pro) 作者 | 劉俐杉 編輯 | 許麗思 聽說機器人甚至已經學會刺繡了? 如果讓機器人搬運重物、抓取零件,你可能早已見怪不怪,讓它搭積木、夾豆子、擰瓶蓋,也已經不算什麼新鮮事了。 但今天上午,臨界點發布的一段視頻,讓靈巧手的“手上功夫”又往前了一步。 鏡頭裡,靈巧手捏著一枚繡花針,將細線穿過針孔,隨著針尖的起落,手指在絲線中穿梭。直至鏡頭拉近,你還會看到靈巧手面對一幅蘇繡,與蘇繡老師配合完成合繡。更加值得注意的是,整段視頻均為實拍,沒有加入任何CG動畫。

5 小時前

驚悚又硬核:研究團隊造出能用指尖行走的機械手

作者:遠洋 責編:遠洋 評論: 9 月 29 日消息,蘇黎世聯邦理工學院(ETH Zurich)軟體機器人實驗室的研究人員製作出了一款可以依靠指尖自行行走的機械手。據瞭解,這款機械手基於市面上一款常規的仿人機械手改造而成,共有五根手指,配備 20 個驅動關節,每根手指擁有四個關節。

20 小時前

海爾洗空氣空調AI科技將巴馬好空氣搬回家

為了一口好空氣,有人給家裡配齊空氣淨化器、新風系統,層層升級呼吸裝備;更有人遠赴廣西巴馬長壽村,斥資置業長居,天不亮就到百魔洞口排隊吸氧。“巴馬好空氣”的說法流傳多年,卻少有人說清它到底好在哪。 9月28日,海爾空調聯合科普博主無窮小亮,深入巴馬長壽村實地探訪,通過實地觀察與專業檢測,拆解巴馬好空氣的真實密碼,也給出了把自然好空氣帶回家的解決方案。 實地溯源:巴馬好空氣不是玄學,是自然的饋贈 每年換季時節,大批外地人從全國各地奔赴巴馬,有人連續多年往返旅居,有人乾脆定居長住,核心訴求都指向當地的空氣。

1 天前

碳硅道統050.5公理母本定稿 構建人機碳硅共生可溯源研究基線

近日,碳硅道統研究體系完成重要迭代工作,050.5公理母本正式凍結定稿並完成歸檔留存。作為001-050卷宗序列的頂層基準文檔,050.5不再新增現象類思辨觀點,旨在為人機碳硅共生方向搭建一套可溯源、可引用、可審計、可復現的底層公理體系,補齊現階段人工智能領域“觀點豐富、公理框架稀缺,論述較多、統一基線不足”的現狀。

2 天前