Reka 發表 Rho-1:一款 19B 參數的全能推理模型,能理解、生成影片並輸出機器人動作

2026年10月6日 06:45
站內 AI 整理稿

Reka has released a research preview of Rho-1, a 19B omni-reasoning model trained from scratch.A single neural network understands and generates text, images and video, reasons over them, and outputs robot actions.

Reka frames it as a direct replacement for agentic pipelines that pass work between modality-specific models.What Rho-1 Changes Most multimodal systems today are pipelines.A central model plans, then hands jobs to specialists for images, video or detection.

Each handoff adds latency, and each specialist sees only a narrow request.Rho-1 removes those handoffs.Text, vision and robotic actions become tokens inside one context window.According to Reka’s research, one unedited session shows the full loop.

The model draws a lighthouse, boxes it, animates it, edits the video into a snowstorm, and explains the difference.That all happens in 5 turns, with no tool call and no second model.

Architecture: Two Streams, One KV Cache Every input and output uses one of two native formats: Discrete tokens carry text, symbolic reasoning and high-level commands.Continuous tokens carry image latents, video frames, robot actions and proprioception.

Each transformer block holds two expert weight streams.The understanding stream handles language and visual parsing.The generation stream denoises latents into images and video.Both streams share attention and operate over the same KV cache.

When a reply needs pixels, the understanding stream emits a discrete handoff token.The generation stream then renders from the full accumulated state.Training combines next-token prediction for discrete sequences with flow matching for continuous outputs.This design has practical effects.

Bounding boxes come out as coordinate tokens, not from a separate detector.A video’s first frame reuses the in-context image representation instead of a re-encoded copy.(function(){window.addEventListener("message",function(e){if(e.data&&e.data.rho1h){var f=document.

getElementById("rho1x-frame");if(f){f.style.height=e.data.rho1h+"px";}}});})(); Speed: Base vs Distilled The base model generates video at 0.79x real-time (median), with a watchable stream starting in roughly 6 seconds.Reka team measured 7.0 seconds to a first clip, against an illustrative 13.

8 seconds for a multi-agent pipeline.A distilled variant cuts denoising from 99 steps to 8, with minimal quality loss reported.It returned a 5.3-second clip in about a second.In Reka’s internal tests, it matched the fastest dedicated image models.

It was also the quickest model tested to the first text token.These are vendor-run tests, not independent benchmarks.World Model and Robotics Rho-1 streams continuously, clip after clip.New instructions enter through the understanding stream and update state mid-rollout.

Reka demonstrates one opening forked into ‘bank left’ and ‘bank right’ continuations.For robotics, actions and future frames decode from the same latent state.A LIBERO simulation episode shows Rho-1 emitting 7 action channels.

To scale past scarce teleoperation logs, Reka pairs Rho-1 with its Inverse Dynamics Model, which infers control signals from raw video.How Rho-1 Compares Data verified on October 5, 2026 from official sources.FeatureReka Rho-1ByteDance BAGELBAAI Emu3.

5Google Genie 3DeveloperRekaByteDance SeedBAAIGoogle DeepMindParameters19B14B total, 7B active (MoT)34BNot disclosedInputsText, image, video, actions, proprioceptionText, imageInterleaved text and imageText prompt, navigation inputsOutputsText, image, video, actions, proprioceptionText, imageInterleaved text and imageInteractive video worldNative video generationYes, capped at 672×384NoNo (image frames, not native clips)Yes, 720p at 24 fpsReal-time steeringYes, continuous rolloutsNoNoYes, promptable world eventsRobot actionsNative continuous action tokensNoEmbodied manipulation demosTakes navigation actions, does not emit themOpen weightsNoYes, Apache 2.

0Yes, Apache 2.0NoAccess todayResearch preview via [email protected] Face, GitHubHugging Face, emu.world appProject Genie for Google AI Ultra subscribersSource/ResourcesReka blogGitHubHugging FaceDeepMind blog Genie 3 access per Project Genie; Emu3.

5 parameter count per its Hugging Face model card.Key Takeaways Rho-1 is a 19B model that reads and writes text, image, video and robot actions.Two expert streams share attention and one KV cache in every block.Distillation cuts denoising from 99 to 8 steps, about 1 second per 5.3 s clip.

Trained on 320 H100s for 3 months; video is capped at 672×384.Research preview only: no public weights, API or pricing yet.Check out the Technical details and the announcement on X.All credit goes to the researcher of this project.

Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?

Connect with us The post Reka Releases Rho-1: A 19B Omni-Reasoning Model That Understands, Generates Video and Outputs Robot Actions in One appeared first on MarkTechPost.

Related

相關文章

Hugging Face Blog生成式AI

Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance

Back to Articles Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance Team Article Published October 6, 2026 Upvote 7 +1 Shaikha Alsuwaidi Shaikha710 Follow tiiuae Omar saif alkaabi Omar-Alkaabi Follow tiiuae Maitha Alhammadi MaithaAlhammadi Follow tiiuae Ahmed Alzubaidi amztheory Follow tiiuae Mohammed Alyafeai Alyafeai Follow tiiuae Leen AlQadi LeenAlQadi Follow tiiuae Basma Boussaha basma-b Follow tiiuae Hakim Hacid HakimHacid Follow tiiuae Arabic is really a family of languages living under one name.

2 小時前
鈦媒體生成式AI

快手的視頻Agent,會不會來晚了?

AI價值官2026.10.05 17:36 · 來自浙江全文4219字00:00 / 11:40視頻模型廠商,正集體從"生成一段畫面"走向"交付一部成片"。文 | AI價值官,作者丨納瓦,編輯丨星野9 月 28 日,快手在模型與 Agent 兩層同時出手:白天,Agent 創作工具 likli 開啟內測;晚間,新一代模型 Kling 4.0 官宣將於 10 月上線。這並非快手的獨門動作,字節、MiniMax、阿里千問都已推出各自的創作 Agent。當模型能力被逐漸拉平,競爭正從模型轉向工作流與商業化。

15 小時前
鈦媒體生成式AI

AI耳機蓄勢,芯片廠商待發

半導體產業縱橫2026.10.05 15:32 · 來自內蒙古全文5079字00:00 / 15:12AI耳機還沒有成為一個邊界清晰的品類,芯片廠商卻已經開始為它準備下一代平臺。文 | 半導體產業縱橫傳統耳機,卒?AI耳機正在無線耳機市場中成為主流。據Research and Markets的測算,2025年全球AI耳機市場規模約為59.9億美元,預計到2026年將攀升至74.2億美元,並在2030年達到173.4億美元。消費電子行業裡最常被重複的一句判斷是:所有硬件,都值得用AI重新做一遍。

17 小時前
鈦媒體生成式AI

硅谷AI,正在開源

影子備忘錄2026.10.05 15:32 · 來自廣東全文4084字00:00 / 10:59大模型從哪家強到哪家合適的轉變。文 | 影子備忘錄過去兩年,硅谷AI圈最常被問到的問題只有一個:哪家的模型最強?但2026年,這個問題正在被另一個更務實的問題取代:完成同一項任務,到底有沒有必要用最貴的模型?今年6月,打車巨頭Uber透露了一個讓整個行業沉默的數字:僅2026年前四個月,公司就耗盡了全年的AI預算,原因僅僅是員工大量使用AI編碼工具。

17 小時前