Dyna Robotics 推出 Dyna-2:以百萬小時人類影片預先訓練的世界動作模型

2026年8月13日 07:42
站內 AI 整理稿

Dyna Robotics has released Dyna-2, a world-action model for robot manipulation.It was pre-trained on more than one million hours of egocentric human video.That is roughly 170 years of continuous waking experience.

Robot learning has been bottlenecked by action-labelled data, which teleoperation must deliberately produce.Dyna-2 tests whether ordinary human video can substitute.The research team trained a data ladder from 1,000 to 1,000,000 hours and measured what scales.

Three results follow: a scaling law on human data, the first transfer of that law to unseen robot data, and evidence that video prediction drives the transfer.Is it deployable?Yes, but as a vendor-operated system, not as downloadable weights.

Dyna Robotics has announced no public checkpoint, API, or license for Dyna-2.Deployment today means buying a Dyna robot cell, not self-hosting a model.Which companies: Dyna-1 robots already run in production in hotels, restaurants, and laundromats, per the company’s August 10, 2026 announcement.

That points at mid-market service operators and multi-site enterprises with repetitive, stationary manipulation work.It is not a fit for solo builders or research labs wanting local inference.

Industries: Hospitality, commercial laundry, food service, light assembly and kitting, and facilities cleaning.

Applications: The 14 post-training tasks map cleanly to real work: trash tray clearing, first-aid kitting, tote construction, food scooping, rope tying, hanger preparation, and targeted drink retrieval from a fridge.

What is Dyna-2 Dyna-2 is a world-action model (WAM): one generative model that denoises future video and a future action chunk, jointly or separately, on a video-diffusion backbone.

It was pre-trained on more than one million hours of egocentric human video, roughly 170 years of continuous waking experience.Architecturally it is a mixture of transformers.Video and action are tokenized separately and get distinct DiT layer stacks that attend to each other.

Proprioception feeds directly into the action transformer.Video tokens use causal masking; action tokens use bidirectional self-attention and attend to context video tokens.Video tokens cross-attend to text, but text does not directly influence action tokens.Training uses flow matching.

A video loss and an action loss share a trunk as two separate marginal velocity fields.Because the action network never takes the noised video latent as an argument, the policy stays reactive at inference — it neither generates nor attends to predicted future video.

The action transformer is deliberately shallower and joins the video stream early, which the team says improves real-time latency without costing performance.

The three scaling results Dyna Robotics cut nested subsets of exactly 1,000, 10,000, 100,000, and 1,000,000 hours, keeping identical proportions from each source.A larger budget only adds data, so curve differences cannot be attributed to distribution shift.

A fixed, disjoint 100-hour validation set scores every rung.A scaling law holds on human data to one million hours: All four metrics improve monotonically and fit power laws: held-out MSE = 0.0691·D^-0.0184 (R²=0.919), [email protected] = 0.357·D^+0.0203 (R²=0.865).Across the ladder, accuracy@0.

1 rises 51% against 12% for MSE.That law transfers to robot data the model never saw: The same checkpoints were scored zero-shot on 39 tasks across two stationary bimanual YAM platforms — 12 internal, 27 from xdof ABC.Zero-shot action MSE = 0.306·D^-0.0713 (R²=0.884).

The team reports an inflection between 10k and 100k hours.The objective matters, and video is a separate axis: Joint denoising beat action-only on 39 of 39 tasks at every action scale.Holding action-labelled data fixed at 50,000 hours and adding video-only hours drops zero-shot robot MSE from 0.

340 to 0.120.Notably, held-out human error does not improve — the benefit of video is specifically cross-embodiment generalization.

On-robot results Each rung was post-trained on 14 tasks, at most 10 hours of robot data each, across three embodiments: 6-DOF YAM arms with parallel-jaw grippers, the same arms with WUJI-2 20-DOF dexterous hands, and a semi-humanoid prototype.

Post-training used robot data only — no human-robot alignment, no co-training.Mean normalized score rose 20% → 28% → 45% → 53% across the ladder, best on 9 of 14 tasks at one million hours.Lockbox Key Turning is the threshold case: 0% up to 100,000 hours, then 90%.

Bottle Cap Untwisting was post-trained on roughly 10 minutes of demonstrations and still climbed to 50%.Against Dyna-1 — the company’s production VLA initialized from Qwen3-VL-4B — an early Dyna-2 reached 1.55× success rate and 1.12× grade, pooled over 7 tasks and 3 checkpoints.

At unseen customer sites, Dyna-2 passed production criteria 87% versus Dyna-1’s 46%, though both pass near 100% in house.A distillation pipeline also cuts video sampling from 10,203 ms to 110 ms on one H100.Interactive explainer ANIMATE THE LADDER Held-out human data · rung 1k hours

Held-out MSE ↓

0.062 D-0.0184 · R²=0.919 Held-out L1 ↓

0.140 D-0.0132 · R²=0.879 Accuracy @0.1 ↑

0.017 D+0.0606 · R²=0.926 Accuracy @0.5 ↑

0.40 D+0.0203 · R²=0.865 Across the full ladder [email protected] rises 51% while MSE improves 12%.Tight-threshold accuracy — a proxy for movement precision — is where scale pays most.<!

-- 03 HUMAN TO ROBOT --> The same checkpoints were scored on 39 robot tasks across two stationary bimanual YAM platforms — 12 internal, 27 from xdof ABC — with zero robot trajectories in pre-training.

Then each rung was post-trained on 14 tasks, at most 10 hours of robot data each, robot data only, no human–robot alignment.ANIMATE THE LADDER Zero-shot on unseen robot data · rung 1k&l

Related

相關文章

WRC 2026|原生全模態世界模型:從模擬世界到交互世界

世界機器人大會期間,智象未來創辦人梅濤於「物理AI引領者論壇」發表演講,提出原生全模態世界模型從「模擬世界」走向「交互世界」的觀點。他強調即使AI模型智商接近140,高IQ不代表全能,需具備在真實物理世界中穩定完成任務的能力,此為Physical AI發展的關鍵。論壇聚焦通用物理智慧的技術演進與產業路徑,匯聚眾多專家參與。

剛剛

阿里巴巴達摩院推出肝癌 AI 模型:可精準識別 1 釐米微小腫瘤

作者:遠洋 責編:遠洋 評論: 感謝網友 HH_KK 的線索投遞!8 月 24 日消息,阿里巴巴達摩院聯合中國醫科大學附屬盛京醫院等機構研發出肝癌診斷 AI 模型 DAMO LiON,可通過 CT 影像識別微小的肝臟癌變病灶。在兩個月的真實世界前瞻臨床試驗中,該 AI 模型發現了 15 例原本被遺漏的惡性腫瘤,絕大部分為 1 釐米左右的病灶,幫助患者得到及時的手術或藥物治療。

剛剛
何夕2077研究與前沿

棋類模型可解釋

在人工智慧研究領域,模型的可解釋性一直是備受關注的課題。近期有觀點指出,棋類模型具備可解釋的特性,這意味著此類模型的決策過程與內部運作機制,能夠被研究者或使用者以相對直觀的方式理解與分析。相較於許多深度學習模型常被視為「黑箱」,棋類模型在處理圍棋、象棋等棋類遊戲時,其每一步的選擇與策略推演,往往能透過棋譜或演算法邏輯加以回溯,從而為AI的透明化提供了一個具體的觀察窗口。

8 小時前

美國專家示警:學生依賴“AI 代寫”會削弱思考能力

作者:清源 責編:清源 評論: 8 月 23 日消息,美國學生使用 AI 完成作業、甚至代寫整篇論文的現象已經十分普遍,也有不少學校允許學生在一定範圍內藉助 AI 工具。據《紐約時報》當地時間 17 日報道,越來越多專家擔心,問題可能不只是學生會不會寫文章,而是長期依賴 AI 可能削弱他們本身的思考能力。

12 小時前