利用 NVIDIA Cosmos3-DROID 建構串流機器人學習管線
In this tutorial, we design an end-to-end streaming robotics learning pipeline around the NVIDIA Cosmos3-DROID dataset without downloading its 707 GB repository locally.We first introspect the LeRobotDataset v3.0 structure and construct a metadata graph from info.
json, task metadata, episode tables, and dataset statistics, then use HTTP byte-range access with PyArrow to selectively read Parquet row groups and columns.
We convert individual episodes into state-action trajectories and analyze joint motion, gripper events, Cartesian end-effector paths, and action-frequency spectra before decoding only the required AV1 video windows through seek-based PyAV/FFmpeg access.
We then normalize observations and actions using dataset statistics, construct an ACT-style chunked PyTorch dataset with optional visual conditioning, and train a multimodal behavior-cloning policy.
Finally, we evaluate the learned policy through open-loop rollout with temporally ensembled action chunks, report per-joint MSE and R^2 against a mean-action baseline, visualize predicted versus ground-truth actions, and save the complete policy checkpoint for downstream use.
Copy CodeCopiedUse a different Browserimport subprocess, sys, os, json, math, time, warnings, random, tempfile warnings.filterwarnings("ignore") subprocess.run([sys.executable, "-m", "pip", "install", "-q", "huggingfacehub>=0.34.0", "pyarrow>=15.0", "av>=12.
0", "pandas", "matplotlib", "tqdm"], check=False) import numpy as np, pandas as pd, pyarrow as pa, pyarrow.parquet as pq import matplotlib.pyplot as plt from huggingfacehub import HfApi, HfFileSystem, hfhubdownload, hfhuburl import torch, torch.nn as nn, torch.nn.functional as F from torch.utils.
data import Dataset, DataLoader REPOID = "nvidia/Cosmos3-DROID" ROOT = "success" VIDEOKEY = "observation.image.wristimageleft" FPS = 15 NEPISODES = 48 HORIZON = 8 OBSHISTORY = 2 USEVISION = True NVISEPS = 6 VISSIZE = 96 EPOCHS = 12 BATCH = 256 SEED = 0 random.seed(SEED); np.random.seed(SEED); torch.
manualseed(SEED) DEV = "cuda" if torch.cuda.isavailable() else "cpu" print(f"[env] torch={torch.version} device={DEV}") if os.environ.get("HFTOKEN"): from huggingfacehub import login; login(os.
environ["HFTOKEN"]) api = HfApi() fs = HfFileSystem() HFS = lambda rel: f"datasets/{REPOID}/{rel}" URL = lambda rel: hfhuburl(REPOID, rel, repotype="dataset") print("\n" + "="78 + "\n1.REPO INTROSPECTION\n" + "="78) allfiles = api.
listrepofiles(REPOID, repotype="dataset") print(f"total files in repo : {len(allfiles):,}") for prefix in ("success/data", "success/videos", "success/meta", "failure/data", "failure/videos", "failure/meta"): print(f" {prefix:<18} {sum(f.
startswith(prefix) for f in allfiles):>6,} files") datashards = sorted(f for f in allfiles if f.startswith(f"{ROOT}/data/") and f.endswith(".parquet")) vidshards = sorted(f for f in allfiles if f.startswith(f"{ROOT}/videos/{VIDEOKEY}/")) metafiles = sorted(f for f in allfiles if f.
startswith(f"{ROOT}/meta/")) print(f"\n[{ROOT}] data shards={len(datashards)} video shards({VIDEOKEY})={len(vidshards)}") print("first data shard :", datashards[0]) print("first video shard:", vidshards[0]) print("\n" + "="78 + "\n2.METADATA\n" + "="78) info = json.
load(open(hfhubdownload(REPOID, f"{ROOT}/meta/info.json", repotype="dataset"))) print(f"episodes={info.get('totalepisodes'):,} frames={info.get('totalframes'):,} " f"tasks={info.get('totaltasks'):,} fps={info.get('fps')}") print("datapath template :", info.
get("datapath")) print("videopath template:", info.get("videopath")) FEATURES = info["features"] statekeys = sorted(k for k in FEATURES if k.startswith("observation.state")) actionkeys = sorted(k for k in FEATURES if k.startswith("action.
")) videokeys = sorted(k for k in FEATURES if FEATURES[k]["dtype"] == "video") print("\nstate :", [f"{k.split('.')[-1]}{tuple(FEATURES[k]['shape'])}" for k in statekeys]) print("action :", [f"{k.split('.')[-1]}{tuple(FEATURES[k]['shape'])}" for k in actionkeys]) print("video :", videokeys) tdf = pd.
readparquet(hfhubdownload(REPOID, f"{ROOT}/meta/tasks.parquet", repotype="dataset")) tdf = tdf.resetindex() tcol = "task" if "task" in tdf.columns else tdf.columns[0] TASKS = dict(zip(tdf["taskindex"].astype(int), tdf[tcol].
astype(str))) if "taskindex" in tdf \ else {i: str(v) for i, v in enumerate(tdf[tcol])} print(f"\n{len(TASKS):,} task strings.Random sample:") for t in random.sample(list(TASKS.values()), min(8, len(TASKS))): print(" ·", t[:90]) epfiles = [f for f in metafiles if "/episodes/" in f and f.endswith(".
parquet")] eps = pd.concat([pd.readparquet(hfhubdownload(REPOID, f, repotype="dataset")) for f in epfiles[:4]], ignoreindex=True) print(f"\nepisodes table: {len(eps):,} rows") print("columns:", [c for c in eps.columns if not c.startswith("stats")][:14], "...
") print(eps[[c for c in ("episodeindex", "length", "data/chunkindex", "data/fileindex") if c in eps.columns]].head()) We initialize the Colab environment, install the required libraries, and configure the Cosmos3-DROID dataset, episode, video, and training parameters.
We inspect the repository structure and identify the available data, video, and metadata shards without downloading the complete dataset.We then load the core metadata and task descriptions to understand the dataset schema, available state/action features, and episode organization.
Copy CodeCopiedUse a different Browserprint("\n" + "="78 + "\n3.BYTE-RANGE PARQUET READER\n" + "="78) def openpf(relpath): return pq.ParquetFile(fs.open(HFS(relpath), "rb")) def rowgroupspan(pf): md, starts, c = pf.metadata, [], 0 for i in range(md.numrowgroups): starts.append(c); c += md.
rowgroup(i).numrows return np.array(starts), c def readrows(pf, lo, hi, columns): starts, total = rowgroupspan(pf) ends = np.append(starts[1:], total) rgs = [i for i in range(len(starts)) if starts[i] < hi and ends[i] > lo] tbl = pf.readrowgroups(rgs, columns=columns) return tbl.
slice(lo - starts[rgs[0]], hi - lo) def col2np(tbl, name): ca = tbl.column(name).combinechunks() if pa.types.islist(ca.type) or pa.types.islargelist(ca.type) or pa.types.isfixedsizelist(ca.type): flat = np.asarray(ca.flatten().tonumpy(zerocopyonly=False)) return flat.reshape(len(ca), -1).astype(np.
float32) return np.asarray(ca.tonumpy(zerocopyonly=False)).reshape(-1, 1).astype(np.float32) SHARD = datashards[0] pf = openpf(SHARD) md = pf.metadata print(f"shard : {SHARD}") print(f"rows : {md.numrows:,} rowgroups: {md.numrowgroups} " f"compressed: {md.serializedsize/1e6:.
1f} MB footer") print(f"columns : {len(pf.schemaarrow.names)}") t0 = time.time() epidxall = pf.read(columns=["episodeindex"]).column("episodeindex").tonumpy() print(f"pulled episodeindex column ({len(epidxall):,} rows) in {time.time()-t0:.1f}s") uniq, firstpos = np.
unique(epidxall, returnindex=True) order = np.argsort(firstpos) uniq = uniq[order]; firstpos = firstpos[order] lastpos = np.
append(firstpos[1:], len(epidxall)) EPBOUNDS = {int(e): (int(a), int(b)) for e, a, b in zip(uniq, firstpos, lastpos)} print(f"{len(EPBOUNDS)} episodes live in this shard " f"(ids {uniq.min()}..{uniq.max()}, mean len {np.mean(lastpos-firstpos):.0f} frames)") STATEUSE = ["observation.state.
jointpositions", "observation.state.gripperposition", "observation.state.cartesianposition"] ACTIONUSE = ["action.jointvelocity", "action.
gripperposition"] READCOLS = STATEUSE + ACTIONUSE + ["timestamp", "frameindex", "taskindex", "episodeindex"] def loadepisode(ep): lo, hi = EPBOUNDS[ep] tbl = readrows(pf, lo, hi, READCOLS) out = {k: col2np(tbl, k) for k in STATEUSE + ACTIONUSE} out["timestamp"] = col2np(tbl, "timestamp").
ravel() out["taskindex"] = int(col2np(tbl, "taskindex").ravel()[0]) out["task"] = TASKS.get(out["taskindex"], "") out["state"] = np.concatenate([out[k] for k in STATE_USE], axis=1) out["action"] = np.concatenate([out[
Related
相關文章

Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade
Most production recommenders are cascades. Candidate generators feed a pre-ranker, which feeds a heavy ranker built on hundreds of engineered features. Yandex’s Sona Technical Report describes a different design.
《Gran Turismo 7》迎來首臺中國VGT 同時GT史上30年首次新增電車駕駛教學
新加坡,2026 年 10 月 3 日 —— 今日,在新加坡 Gran Turismo World Series(GT World Series)賽事現場,Gran Turismo系列遊戲製作人山內一典與小米汽車歐洲研發中心設計負責人Jean-Arthur Madelaine現場聯合宣佈:Xiaomi Vision Gran Turismo(小米 Vision GT)即將於10月正式上線《Gran Turismo 7》(GT7),成為該。
A Coding Guide to Google Research’s Kauldron: Configs That Are Plain Data, Components Wired by String, and a JAX Trainer You Can Read End to End
In this tutorial, we implement Kauldron, the JAX training library from Google Research that describes itself as optimized for research velocity and modularity, and we take those two words literally by testing what they actually buy us.
剛剛,iQOO掏出年度旗艦,自研電競芯片性能提升15%,首款電競平板也來了
作者 | 陳駿達 編輯 | 心緣 9月29日報道,剛剛,vivo旗下iQOO品牌發佈了年度旗艦手機iQOO 16,這臺手機搭載了第六代驍龍8超級至尊版,全球首發了由iQOO和三星顯示聯合研發的2K 165Hz三星珠峰屏,並基於iQOO搭建的“3+2遊戲技術版圖”,提升了手機在視效、操控、直播和跨端遊戲等維度的體驗。
抽“錦鯉”享美食!“點亮杭州 碰見好運”服務消費季活動啟動
本文作者: Nemo 2026-09-25 10:17 導語:據瞭解,圍繞“點亮杭州 碰見好運”主題,活動將在9月21日至10月7日期間發放百萬級消費券,覆蓋吃喝玩樂購。9月24日,“點亮杭州 碰見好運”服務消費季活動在西湖區天目裡國際街區正式啟動。
聚焦院外管理提質增效|《急性冠狀動脈綜合徵患者院外長期隨訪管理共識》更新研討,胸痛中心智慧全程管理行動項目正式啟動
近日,第八屆“儒道心學”心血管病學會議、第十屆滬魯心血管病專家論壇、第九屆日照心血管峰會在山東日照召開。由葛均波院士領銜,黃愷、蘇國海、李春潔等數十位心血管領域權威專家參與,會上完成兩大核心動作:一是召開《急性冠狀動脈綜合徵患者院外長期隨訪管理共識》更新研討會,專家集體錨定共識修訂的核心方向;二是胸痛中心智慧全程管理行動項目正式啟動,以專家共識為指引推進先行落地驗證。作為醫療 AI 賦能院外管理創新的先行者,訊飛醫療執行總裁鹿曉亮受邀參會,與學界、業界共同推動心血管院外管理向標準化、智能化、全週期階段邁進。