τ_0-VLA長程操控
Papers arxiv:2608.
16885 Copy markdown τ0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation Published on Aug 17 · Submitted by jrryzh(SII) on Aug 21 · Shanghai Innovation Institute Upvote 14 +6 Authors: Xiaowei Cai ,Yunuo Cai ,Bingao Chen ,Jingxiao Chen ,Zhi Chen ,Siyuan Feng ,Tengyu Hou ,Jingshun Huang ,Han Jiang ,Runkun Ju ,Dong Li ,Mingxiang Li ,Shaowei Li ,Xinchen Li ,Yifan Li ,Yi Liu ,Zhongyuan Liu ,Jianlan Luo ,Junwen Miao ,Ruiqi Ni ,Buqing Nie ,Mingjie Pan +17 authors Abstract A hierarchical vision-language-action model improves long-horizon robot manipulation by using world-model-guided test-time search to scale computation for high-level subtask decisions.
Generated by thinkingmachines/Inkling-Small Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks.
Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices.
We introduce τ0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation.
At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output.A low-level policy then executes the generated subtask across multiple robot embodiments.
The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training.
Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.
View arXiv page View PDF Project page GitHub 528 Add to collection Community J3rr1 Paper author Paper submitter 2 days ago 🤖 What if a robot could compare possible futures before deciding what to do next?We introduce τ₀-VLA, a hierarchical robot foundation model for long-horizon manipulation.
Its high-level policy maintains execution memory and, when a decision is uncertain, allocates additional test-time computation to propose candidate subtasks, predict their visual consequences with a world model, and compare alternatives before committing.
A generalist low-level VLA then executes the selected subtask across robot embodiments.Highlights: The low-level policy is trained on 40,115 hours of heterogeneous real-world robot data with multimodal co-training.
Selective test-time computation improves next-subtask prediction accuracy by 15–24 percentage points across in-domain and distribution-shifted settings.We evaluate real-world manipulation tasks containing 13–25 ordered steps, with episodes lasting up to 12 minutes.
Using the same low-level policy, hierarchical planning improves average closed-loop success from 27.5% to 45.0% across four long-horizon tasks.We release the official code and pretrained low-level VLA checkpoint, with high-level policy on the way.
🌐 Project page💻 Code🤗 Model checkpoint Questions and feedback are very welcome!Reply librarian-bot 2 days ago This is an automated message from the Librarian Bot.I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation (2026) Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models (2026) WorldScape Policy 2.
0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory (2026) StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models (2026) Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation (2026) HarnessWAM: Bridging Prediction and Deliberation in World Action Models (2026) G0.
5: One Autoregressive Stream for Robot Reasoning and Action (2026) Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend Reply EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images Comment · Sign up or log in to comment Upvote 14 +2 Get this paper in your agent: hf papers read 2608.16885 Don't have the latest CLI?curl -LsSf https://hf.co/cli/install.
sh | bash Models citing this paper 1 Datasets citing this paper 0 No dataset linking this paper Cite arxiv.org/abs/2608.16885 in a dataset README.md to link it from this page.Spaces citing this paper 0 No Space linking this paper Cite arxiv.org/abs/2608.16885 in a Space README.
md to link it from this page.Collections including this paper 4
Related
相關文章

消息稱 NVIDIA 考慮向 Perplexity AI 投資“數十億美元”
作者:溯波(實習) 責編:溯波 評論: 8 月 24 日消息,外媒 The Information 稍早前報道稱,NVIDIA(英偉達)考慮在 Perplexity AI 的最新融資輪中向這家人工智能初創企業投資“數十億美元”。雙方正就該交易展開磋商,還可能達成技術授權協議。

Rabbit做了個“所有Agents的Agent” ,體驗後我覺得Agent早就該長它這樣
Rabbit推出了一個被稱為「所有Agents的Agent」的新產品,實際體驗後讓人覺得Agent早就該長這樣。文章認為Agent普及的關鍵並非增加更多能力,而是讓用戶少理解十個概念,降低使用門檻。
Guidelight評估五大AI實驗室,OpenAI遏制能力排名第一
該評估覆蓋Anthropic、Google、OpenAI、Meta和xAI,僅依據公開信息考察其是否建立監控、異常行為處置、第三方審計及失控模型關閉等機制。在滿分5分的評估中,OpenAI以3分排名最高,Anthropic和Meta得分最低。