Zero-WAM看視頻幹活

2026年8月30日 00:00
站內 AI 整理稿

Papers arxiv:2608.

26103 Copy markdown Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization Published on Aug 26 · Submitted by Yujie Zhao on Aug 28 · Robbyant Research Upvote 17 +9 Authors: Jiaming Zhou ,Qihang Zhang ,Gangwei Xu ,Cunxin Fan ,Yujie Zhao ,Ruilin Wang ,Yiming Luo ,Shuai Yang ,Xing Zhu ,Yujun Shen ,Junwei Liang ,Yinghao Xu Abstract Zero-WAM enables robotic manipulation of unseen tasks by conditioning a causal video-action model on in-context human video guidance, supported by an automatically generated dataset and a future-chunk prediction objective.

Generated by thinkingmachines/Inkling-Small Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning.

In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update.This form of in-context learning (ICL) turns generalization into a problem of task specification.

To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution.

We present Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance.

To address the scarcity of task-rich paired human-robot data, we propose an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks.

For model training, we further introduce an in-context future chunk prediction (IFP) objective that suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt.On seven unseen tasks in RoboTwin 2.0 simulation, Zero-WAM achieves a 47.

0% average success rate, an absolute improvement of 29.5 percentage points over the strongest video-action baseline.

In real-world evaluations, it follows human video guidance to generalize to unseen task configurations involving multi-object scenes, long-horizon manipulation, and fine-grained insertion.

View arXiv page View PDF Project page GitHub 158 Add to collection Community HomieZ Paper submitter 2 days ago Webiste: https://robbyant-research.github.io/Zero-WAM Code: https://github.com/robbyant-research/Zero-WAM Paper: https://arxiv.org/abs/2608.

26103 Reply librarian-bot 1 day ago This is an automated message from the Librarian Bot.I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API Native Video-Action Pretraining for Generalizable Robot Control (2026) WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time (2026) WorldScape Policy 2.

0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory (2026) JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment (2026) DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation (2026) GigaWorld-Policy-0.

5: A Faster and Stronger WAM Empowered by AutoResearch (2026) StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation (2026) Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend Reply EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images Comment · Sign up or log in to comment Upvote 17 +5 Get this paper in your agent: hf papers read 2608.26103 Don't have the latest CLI?curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0 No model linking this paper Cite arxiv.org/abs/2608.

26103 in a model README.md to link it from this page.Datasets citing this paper 0 No dataset linking this paper Cite arxiv.org/abs/2608.26103 in a dataset README.md to link it from this page.Spaces citing this paper 0 No Space linking this paper Cite arxiv.org/abs/2608.26103 in a Space README.

md to link it from this page.Collections including this paper 2

Related

相關文章

OpenAI被曝按噸囤Mac mini,幾萬臺全拿去訓AI了

新智元·2026年08月31日 11:53AI實驗室按噸囤Mac mini訓Agent,帶火蘋果Mac營收。等等,AI實驗室囤Mac mini,竟然是按「噸」囤的?!The Information剛剛爆料,OpenAI已經採購了數萬臺Mac mini和Mac Studio,到現在還在追著蘋果要貨。

剛剛

Grok Bot王炸登場,深度接入X,全自動監測全球最新動態

新智元·2026年08月31日 11:52全自動全球熱點監測最關鍵一環。Grok Bot 現在可以直接接入 X 了。它不只是「能讀 X 了」這麼簡單。你在 Grok Bot 裡關聯你的 X 賬號,系統自動幫你創建開發者賬號,付費用戶還直接送你一筆 X API 的調用額度。

剛剛

Cisco 給 9 萬員工配上個人 Agent:記住你的一切,還能替你跨系統辦事

思科(Cisco)近日在公司內部展開一項大規模的人工智慧部署計畫,一口氣為旗下多達九萬名員工配備專屬的個人 Agent。這款數位助理並非簡單的問答機器人,而是被賦予「記憶」能力,能夠記住每位員工的工作習慣、職務內容與日常需求,更關鍵的是,它具備跨系統執行任務的能力,可以直接代辦許多繁瑣的流程性工作。這項決策在企業軟體與 AI 應用領域投下震撼彈,也讓外界得以一窺大型科技公司如何將生成式 AI 真正落地到內部營運。 根據了解,這套個人 Agent 的核心設計理念在於「理解你的一切」。

剛剛
量子位AI Agent

OpenAI買幾萬臺Mac搞強化訓練!英偉達的活被蘋果搶了

OpenAI和Anthropic大量採購Mac mini與Mac Studio進行強化學習訓練,因其統一內存架構在AI任務上具優勢,帶動Mac季度銷售額成長近29%。蘋果意外在企業AI市場取得成功,但面臨供應短缺與英偉達推出競爭產品DGX Spark的挑戰。

剛剛
量子位AI Agent

全國第三,公司第二,“初創黑馬”靈犀智湧用ROSS Harness把機器人送進工業具身智能第一梯隊

靈犀智湧用一臺由Demo級本體組裝而成的機器人,成為工業場景賽除行業頭部企業外唯一獲獎的機器人公司。 第二屆世界人形機器人運動會的51個賽項,劃出了兩條清晰的評價體系:競技賽檢驗運動極限與協同能力,場景賽則直接對標真實崗位的作業標準。工業場景裝配上料崗正是後者中最貼近產線的賽項,評判標準直指工業級交付的核心指標:長時程穩定性、毫米級精度,以及擾動下的自主恢復能力。

剛剛
何夕2077AI Agent

多智能體做數學

Computer Science > Artificial Intelligence arXiv:2608.23691 (cs) [Submitted on 24 Aug 2026] Title:Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment Authors:Stephen Chung, Wenyu Du, William J.

4 小時前