τ_0-VLA長程操控

2026年8月24日 00:00
站內 AI 整理稿

Papers arxiv:2608.

16885 Copy markdown τ0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation Published on Aug 17 · Submitted by jrryzh(SII) on Aug 21 · Shanghai Innovation Institute Upvote 14 +6 Authors: Xiaowei Cai ,Yunuo Cai ,Bingao Chen ,Jingxiao Chen ,Zhi Chen ,Siyuan Feng ,Tengyu Hou ,Jingshun Huang ,Han Jiang ,Runkun Ju ,Dong Li ,Mingxiang Li ,Shaowei Li ,Xinchen Li ,Yifan Li ,Yi Liu ,Zhongyuan Liu ,Jianlan Luo ,Junwen Miao ,Ruiqi Ni ,Buqing Nie ,Mingjie Pan +17 authors Abstract A hierarchical vision-language-action model improves long-horizon robot manipulation by using world-model-guided test-time search to scale computation for high-level subtask decisions.

Generated by thinkingmachines/Inkling-Small Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks.

Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices.

We introduce τ0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation.

At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output.A low-level policy then executes the generated subtask across multiple robot embodiments.

The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training.

Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.

View arXiv page View PDF Project page GitHub 528 Add to collection Community J3rr1 Paper author Paper submitter 2 days ago 🤖 What if a robot could compare possible futures before deciding what to do next?We introduce τ₀-VLA, a hierarchical robot foundation model for long-horizon manipulation.

Its high-level policy maintains execution memory and, when a decision is uncertain, allocates additional test-time computation to propose candidate subtasks, predict their visual consequences with a world model, and compare alternatives before committing.

A generalist low-level VLA then executes the selected subtask across robot embodiments.Highlights: The low-level policy is trained on 40,115 hours of heterogeneous real-world robot data with multimodal co-training.

Selective test-time computation improves next-subtask prediction accuracy by 15–24 percentage points across in-domain and distribution-shifted settings.We evaluate real-world manipulation tasks containing 13–25 ordered steps, with episodes lasting up to 12 minutes.

Using the same low-level policy, hierarchical planning improves average closed-loop success from 27.5% to 45.0% across four long-horizon tasks.We release the official code and pretrained low-level VLA checkpoint, with high-level policy on the way.

🌐 Project page💻 Code🤗 Model checkpoint Questions and feedback are very welcome!Reply librarian-bot 2 days ago This is an automated message from the Librarian Bot.I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation (2026) Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models (2026) WorldScape Policy 2.

0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory (2026) StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models (2026) Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation (2026) HarnessWAM: Bridging Prediction and Deliberation in World Action Models (2026) G0.

5: One Autoregressive Stream for Robot Reasoning and Action (2026) Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend Reply EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images Comment · Sign up or log in to comment Upvote 14 +2 Get this paper in your agent: hf papers read 2608.16885 Don't have the latest CLI?curl -LsSf https://hf.co/cli/install.

sh | bash Models citing this paper 1 Datasets citing this paper 0 No dataset linking this paper Cite arxiv.org/abs/2608.16885 in a dataset README.md to link it from this page.Spaces citing this paper 0 No Space linking this paper Cite arxiv.org/abs/2608.16885 in a Space README.

md to link it from this page.Collections including this paper 4

Related

相關文章

IT之家AI Agent

消息稱 NVIDIA 考慮向 Perplexity AI 投資“數十億美元”

作者:溯波(實習) 責編:溯波 評論: 8 月 24 日消息,外媒 The Information 稍早前報道稱,NVIDIA(英偉達)考慮在 Perplexity AI 的最新融資輪中向這家人工智能初創企業投資“數十億美元”。雙方正就該交易展開磋商,還可能達成技術授權協議。

剛剛
AIBaseAI Agent

Guidelight評估五大AI實驗室,OpenAI遏制能力排名第一

該評估覆蓋Anthropic、Google、OpenAI、Meta和xAI,僅依據公開信息考察其是否建立監控、異常行為處置、第三方審計及失控模型關閉等機制。在滿分5分的評估中,OpenAI以3分排名最高,Anthropic和Meta得分最低。

8 小時前7400
何夕2077AI Agent

智能體技能合集

VoltAgent 的智能體技能庫近期持續擴充,目前收錄的技能規模已突破千項,相關倉庫在開發社群中獲得約 31.3k 的星標關注。這個持續成長的技能庫,正逐步成為開發者快速組合與部署智能體工作流程的重要資源。 該技能庫涵蓋多種類型的 CLI 工作流程,這些流程具備高度可複用性,讓開發者不必從零開始建構每個環節,而是能直接引用既有技能來加速專案進度。隨著技能項目不斷增加,VoltAgent 生態系的應用範圍也隨之擴大,從自動化任務到複雜的指令處理,都能找到對應的現成模組。

9 小時前
何夕2077AI Agent

ruflo蜂群代理

近期在AI資訊日報中受到矚目的「ruflo蜂群代理」,是一款以多智能體框架為核心的代理管理工具。其設計重點在於讓開發者能夠有效管理複雜的代理流,透過系統化的方式協調多個AI代理協同運作,而非僅是單一任務的自動化處理。這種架構特別適合需要分工、接力或平行處理的應用場景,讓整個執行流程更為清晰且可控。 值得注意的是,ruflo導入RAG(檢索增強生成)記憶機制,讓代理在運作過程中能將經驗與知識「沉澱」下來。也就是說,代理不只是每次從頭開始執行,而是可以從過往的互動或任務中提取相關資訊,作為後續決策的參考依據。

9 小時前
何夕2077AI Agent

調研不是問答

近日在 Reddit 的「研究工作流」討論串中,有使用者發文抱怨現行調研流程的碎片化問題。該用戶指出,儘管市面上有許多 AI 工具能協助蒐集與整理資訊,但多源材料的彙整依然高度依賴人工手動搬運,未能真正打通從資料獲取到最終產出的環節。這篇貼文引發不少共鳴,凸顯出當前調研工作的一大痛點:工具雖多,但流程仍舊零散。 該用戶進一步分析,這類工作流的瓶頸並不在於生成文本的速度或品質,而是如何將零散的資訊整合成具有決策價值的報告。

9 小時前