Generalist AI Releases GEN-1.5: A Robot Foundation Model That Learns New Tasks From One 3–12 Second Demo
Generalist AI has released GEN-1.5, a robot foundation model that learns a new physical task from a single demonstration.Drop 3–12 seconds of sensorimotor data into its 30-second context window, and the robot performs the task.No gradient updates, no fine-tuning, no task-specific programming.
Across 10 diverse manipulation tasks, this one-shot in-context prompting averaged 59% success (±10% std.dev.) straight from the pretrained model.Ten gradient steps on five minutes of data per task raised that to 83% (±9%).
Generalist calls the mechanism physical prompting, and says it was never trained for: no architectural changes, no meta-learning loop, no auxiliary objectives.It emerged from over eight months of continuous pretraining on physical interaction data.
The tasks are simple and short-horizon, and the company says so plainly.But this is the first model its team knows of where one-shot learning of physical skills has emerged at scale.Is it deployable?Not yet — this is a research release.
There are no public weights, no API, no pricing page and no self-serve product.Generalist AI runs GEN-1.5 on its own fleet and data engine.Anyone who wants it today goes through a direct partnership.What is GEN-1.5?GEN-1.
5 is a large multimodal model that takes video, sensor, language and proprioceptive inputs, holds 30 seconds of memory, and emits 100 Hz action trajectories.It has been pretraining continuously for over eight months on physical interaction data captured in homes, warehouses and factories.
The main mechanism is physical prompting.A sensorimotor example — sensor streams plus the action trajectory — is inserted into the 30-second context window through a drag-and-drop interface.The remainder of the window holds rolling observations.
The model then performs the task immediately, with zero gradient steps and no fine-tuning.Crucially, none of this was designed in.Generalist states there were no architectural changes to promote in-context learning, no meta-learning loop, and no auxiliary objectives encouraging improvisation.
The capability emerged from pretraining scale, the same way one-shot prompting emerged in GPT-3.The numbers Across 10 diverse tasks, one-shot in-context prompting averaged 59% success (±10% std.dev.) from the pretrained model, with no training at all.
Ten gradient steps on five minutes of data per task — roughly 50 demonstrations — raised that to 83% (±9%).In the extreme case, one gradient step on one minute of data reached 66.5% on a held-out task, with no adaptation-specific hyperparameter sweep.The compute story is the interesting part.
Adapting robot policies has typically taken tens of thousands of gradient steps.Ten steps here move the model weights on held-out tasks by less than 0.15%, which suggests fine-tuning is reconfiguring knowledge the model already has rather than building new representations.
Generalist frames it as test-time training in an extremely low-data regime.Three transfer results worth knowing Compositional generalization: Two independently recorded prompts placed in context get chained into one continuous behaviour.
The model produces the bridging motions — repositioning, regrasping, error recovery — that appear in neither demonstration.
Zero-shot sim-to-real: A demonstration recorded entirely in simulation works as a prompt for the real robot, despite pretraining containing no simulation data — neither rendered video nor simulated dynamics.For some tasks, demonstrations no longer need to be collected physically.
Human-to-robot imitation: In some cases a person demonstrates with their own hands, in view of the robot’s cameras, and the model reproduces it with the robot’s hands.Generalization also shows up after light fine-tuning.
Trained on five minutes of brushing a block into a bowl, the model used a banana as a makeshift brush, and used a dustpan to lift and dump the block instead — a different contact sequence entirely.
It also removed a sheet of paper covering the bowl, and worked ambidextrously when demonstrations used one hand.Interactive explainer #mtp-gen15-embed{background:#0A0A0A !important;border:1px solid #2A2A2A !important;border-radius:14px !important;overflow:hidden !important;margin:26px 0 !
important;padding:0 !important;color:#F2F2F2 !important} #mtp-gen15-embed iframe{display:block !important;width:100% !important;border:0 !important;background:#0A0A0A !important;min-height:600px !
important} #mtp-gen15-embed hr,#mtp-gen15-embed p:empty,#mtp-gen15-embed del,#mtp-gen15-embed s{display:none !important} #mtp-gen15-embed p{margin:0 !important;padding:0 !important;height:1px !important} #mtp-gen15-embed pre,#mtp-gen15-embed code{background:#131313 !important;color:#F2F2F2 !
important;border:1px solid #2A2A2A !important} @media (max-width:640px){#mtp-gen15-embed iframe{min-height:640px !important}} (function(){ window.addEventListener('message',function(e){ var d=e.data; if(d&&d.mtpFrame==='gen15'&&d.height){ var f=document.getElementById('mtp-gen15-frame'); if(f){f.
style.height=d.height+'px';f.setAttribute('height',d.height);} } },false); })(); Key Takeaways GEN-1.5 learns new manipulation tasks from a single 3–12 second demonstration dropped into its 30-second context window.
One-shot in-context prompting hit 59% across 10 tasks; 10 gradient steps on 5 minutes of data hit 83%.One-shot, sim-to-real and human-to-robot transfer emerged from pretraining — none of it was explicitly trained for.Ten gradient steps change weights by under 0.
15%, collapsing per-task adaptation compute by orders of magnitude.No weights, no API, no product: treat this as a research signal about scaling, not a deployable system.Check out the GEN-1.5 research post and @GeneralistAI announcement thread.
Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Generalist AI Releases GEN-1.5: A Robot Foundation Model That Learns New Tasks From One 3–12 Second Demo appeared first on MarkTechPost.
Related
相關文章

1199 元:小米米家立式學習燈 2C 上架,明日開啟預約
這款新品主打長頭燈設計,採用雙光源結構,最高亮度可達 11000lm,通過多項權威安全認證,支持米家 App 互聯與雷達感應自動開關燈,提供 3 年質保。#小米米家新品#

AI偶像,不能照搬真人明星的邏輯
虛眸2026.08.24 14:26 · 來自北京全文4075字00:00 / 11:50就算看著再逼真,也知道不是人。文 | 虛眸第一波AI明星出道,並不順利。AI短劇《被裁掉的女孩》的虛擬女主角方桃子,代言隱形眼鏡時稱“戴了一天很舒服”,隨即被大眾質疑:一個沒有身體的AI角色,如何感受“舒服”?《與你深情,侵入餘生》中的男女主段宴和容寄僑,以演員身份二搭“出演”新劇《分手後男頻女頻大亂鬥》,讓粉絲對自家偶像究竟是誰、屬於哪個次元的世界產生認知混亂。

小米米家智能魚缸 2 Pro 開啟眾籌:支持自動餵食,眾籌價 599 元
米家智能魚缸 2 Pro 今日在小米有品開啟眾籌,售價 599 元。產品配備定製循環水路、自動餵食器及 1.47 寸 LCD 彩屏,支持米家 App 遠程操控與小愛同學語音控制,還可根據魚種信息智能匹配運行策略。#小米有品# 你心動了嗎?
泰康重磅發佈養醫大模型1.0,全面重塑全生命週期醫養服務
泰康保險集團近日正式發佈自研養醫垂類大模型1.0版本,依託其醫養康寧無縫對接服務體系、長壽醫療學科佈局及長壽隊列獨家數據積累,正探索具有自身特色的差異化垂類大模型發展道路。總裁劉挺軍表示,發佈不是終點而是全新起點;該模型是泰康新壽險在科技維度落地“全生命週期醫養康寧”的重要實踐。
Twitch 遭千名主播集體起訴:亞馬遜默認用直播內容"喂"AI,退出還不具追溯力
Twitch因未經同意、未付酬將主播直播、剪輯、聊天記錄等用於訓練亞馬遜生成式AI,遭集體訴訟。原告已提交37頁訴狀,指控平臺違反默示協議及不公平競爭法。爭議焦點是AI訓練默認開啟、退出機制形同虛設,引發創作者社區強烈不滿。