MarkTechPost AI生成式AI

Induction Labs Photon-1 Simulates Desktops, Plays Checkers, and Models Billiard Physics From One Pretraining Run

2026年7月26日 09:14

重點摘要

Most agents that learn from video need to know what action produced each frame. Induction Labs is arguing that this requirement is the bottleneck. Last week, they released imagination models, a foundation model architecture that pretrains on raw video with no action labels at all. Their test system is Photon-1, a sparse 106B-A5B mixture-of-experts (MoE) transformer trained on 18 years of computer demonstration video. On an internal computer use benchmark, Induction Labs reports that Photon-1 beats Gemini 3.1 Flash-Lite while using far less pretraining compute and costing roughly 3× less to serve. What an imagination model actually does An imagination model predicts future frames autoregressively using a next-latent-token-prediction objective. It does not generate pixels during pretraining.

站內 AI 整理稿

Most agents that learn from video need to know what action produced each frame. Induction Labs is arguing that this requirement is the bottleneck. Last week, they released imagination models, a foundation model architecture that pretrains on raw video with no action labels at all. Their test system is Photon-1, a sparse 106B-A5B mixture-of-experts (MoE) transformer trained on 18 years of computer demonstration video. On an internal computer use benchmark, Induction Labs reports that Photon-1 beats Gemini 3.1 Flash-Lite while using far less pretraining compute and costing roughly 3× less to serve. What an imagination model actually does An imagination model predicts future frames autoregressively using a next-latent-token-prediction objective. It does not generate pixels during pretraining. Everything is modeled in a learned representation space. The claim that matters is this: predicting future states teaches the model to complete tasks, even though it never sees an action during pretraining. Induction Labs calls this an implicit policy. The model learns concepts of what a person is doing, rather than a label for each mouse click. The compression trick that makes it scale The architecture depends on a vision encoder that uses finite scalar quantization (FSQ). Each frame is compressed into 960 discrete tokens. Each token is an 8-dimensional vector. Each dimension takes one of five values: −1, −1/2, 0, 1/2, 1. That gives a codebook of 5⁸ possible codes. The resulting encoding is about 2.2 KB per frame. Induction Labs reports over 100× better compression than existing OCR and multimodal-model representations, while preserving text, layout and state changes. To hit that rate, Photon-1 uses a differential latent encoder. It encodes video frames as pairs, so the latents describe differences between frames rather than frame contents. (function(){ window.addEventListener('message',function(e){ if(!e.data||e.data.mtpFrame!=='mtp-photon-explainer')return; var f=document.getElementById('mtp-photon-explainer-frame'); if(f&&e.data.height)f.style.height=e.data.height+'px'; }); })(); Data and pretraining compute The corpus starts from an internal index of 2 billion publicly available videos. Filtering reduces that to roughly 2 million computer screen recordings. An internal keyframe detection model strips redundant frames. The final dataset is 575 million frames, sampled at 1 frame per second. That equals 552 billion tokens, or about 18 years of video. Photon-1 was pretrained from scratch for a single epoch. Training the 106B-A5B MoE at 32K context took approximately 30,000 H200 GPU-hours, or 4.4×10²² training FLOPs. The research team implemented training in PyTorch with custom fused kernels for the vision encoder and MoE layers, sustaining 40% end-to-end MFU. Those three figures are mutually consistent: 30,000 H200-hours at 40% MFU lands almost exactly on 4.3×10²². From imagination to action Induction Labs finetuned Photon-1 on fewer than 35,000 computer use trajectories to teach the action and instruction format. Special computer use tokens let the model emit actions. At inference, Photon-1 predicts the next frame’s state first, then outputs the action that gets there. Online reinforcement learning follows. Rollouts run in real time on virtual machines at scale, and outcomes are verified programmatically to produce reward. The Linux VMs run five desktop environments (LXQt, Xfce, MATE, GNOME and Plasma), each with a Google account for login-restricted web apps and an internal ChatGPT clone with no rate limits. The compute and cost comparison ModelPretraining computeWeighted inference cost / 1M tokens*Gemini 3.1 Flash-Lite1.200 × 10²⁴ FLOPs$0.36Photon-10.044 × 10²⁴ FLOPs$0.11 *Weighted at a 10:1 input-to-output token ratio, which Induction Labs says matches its computer use tests. Two caveats belong next to that table. First, the Gemini figure is Induction Labs’ own conservative estimate, assuming 8B active parameters and 25T pretraining tokens. Taken at face value the ratio is about 27×, not the 30× headline; Induction Labs states “at least 30×” on the basis that the true Gemini number is likely higher and the model was likely distilled. Second, the benchmark is internal and unreleased, so the result is not independently reproducible today. Photon-1’s own breakeven cost on Induction Labs’ hardware is $0.06 per 1M input tokens and $0.60 per 1M output tokens, with no speculative decoding. Does it generalize past the desktop? This is the more interesting test, because Photon-1 saw only computer use video. The research team finetuned it on domains absent from pretraining and compared against two baselines: a vision encoder baseline with the same architecture and size but no imagination pretraining, and an LLM baseline (Ling-flash-2.0 from Inclusion AI, pretrained on 20T tokens). On 20,000 tournament checkers games from the Open Checkers Archive 2.0, Photon-1 beat both baselines on world simulation and on move quality. On 10,000 synthetically generated billiard games simulated at 5 fps, it produced a mean absolute error of 0.47 against the ground-truth physics engine, versus 1.15 for the LLM baseline and 1.44 for the vision encoder baseline. Photon-1 also picked up human priors from the pretraining video. After RL, it learned to use the in-VM ChatGPT clone to draft artifacts and answer knowledge questions, steering the LLM the way a person would. Key Takeaways Photon-1 learns an implicit policy from 18 years of screen recordings with zero action labels, using next-latent-token prediction. FSQ compresses each frame to 960 tokens (~2.2 KB), a reported 100× gain over OCR and multimodal representations. Trained for ~30,000 H200 GPU-hours, it beats Gemini 3.1 Flash-Lite on an internal benchmark at ~27× less pretraining compute. Despite seeing only desktop video, it beats an LLM baseline at checkers and billiard physics after finetuning. No weights, no API, no license — this is a research result, not a deployable model. Check out the full technical writeup from Induction Labs and the announcement thread on X. All credit for this research goes to the researchers of this project. The post Induction Labs Photon-1 Simulates Desktops, Plays Checkers, and Models Billiard Physics From One Pretraining Run appeared first on MarkTechPost.

Related

相關文章

OpenAI 和 Anthropic,正在砸錢搶這個市場

賬號設置我的關注我的收藏申請的報道退出登錄登錄搜索36氪Auto數字時氪未來消費智能湧現未來城市啟動Power on36氪出海36氪研究院潮生TIDE36氪企服點評36氪財經職場bonus36碳後浪研究所暗湧Waves硬氪氪睿研究院媒體品牌企業號企服點評36Kr研究院36Kr創新諮詢企業服務核心服務城市之窗政府服務創投發佈LP源計劃VClubVClub投資機。

剛剛

三分之一arXiv淪陷,CS論文65%被判“AI味”,數學僅0.7%

賬號設置我的關注我的收藏申請的報道退出登錄登錄搜索36氪Auto數字時氪未來消費智能湧現未來城市啟動Power on36氪出海36氪研究院潮生TIDE36氪企服點評36氪財經職場bonus36碳後浪研究所暗湧Waves硬氪氪睿研究院媒體品牌企業號企服點評36Kr研究院36Kr創新諮詢企業服務核心服務城市之窗政府服務創投發佈LP源計劃VClubVClub投資機。

剛剛

決戰8月,Anthropic甩出Fable 5.1,奧特曼攜GPT-6連夜彙報

賬號設置我的關注我的收藏申請的報道退出登錄登錄搜索36氪Auto數字時氪未來消費智能湧現未來城市啟動Power on36氪出海36氪研究院潮生TIDE36氪企服點評36氪財經職場bonus36碳後浪研究所暗湧Waves硬氪氪睿研究院媒體品牌企業號企服點評36Kr研究院36Kr創新諮詢企業服務核心服務城市之窗政府服務創投發佈LP源計劃VClubVClub投資機構庫投資機構職位推介投資人認證投資人服務尋求報道36氪Pro創投氪堂企業入駐創業者服務創投平臺AI測評網 首頁快訊資訊推薦財經AI自助報道浙江最新創投汽車科技專精特新直播視頻專題活動搜索尋求報道我要入駐城市合作決戰8月,Anthropic甩出Fable 5.1,奧特曼攜GPT-6連夜彙報新智元·2026年07月27日 16:01八月的模型戰場,硝煙已經點燃 GPT-6和Fable 5.1,已經準備就位,很可能8月上線。 據悉,本週奧特曼已經突降華盛頓,展示了OpenAI有史以來的最強大模型——GPT-6。 據外媒Axios爆料,GPT-6 已經具備了進行原創科學研究的能力,並在內部測試中通過不休不眠的智能體集群(Agent Swarms),展現出危險的長程規劃與自主滲透能力。 而就在同一時刻,Anthropic這邊也有消息洩露:Fable 5.1 已經就位,瞄準 8 月,加量不加價,試圖以田忌賽馬戰術,在GPT-6落地前完成狙擊。 這個8月,註定不平靜,一場肉搏戰即將打響,目標直指ASI的奇點時刻。 Anthropic的暗中狙擊:Fable 5.1已準備好,等待一擊斃命 Anthropic 顯然已經嗅到奇點逼近的氣息。 業內爆料證實,Anthropic 籌備已久的 Fable 5.1 已經完成內部測試,瞄準 8 月發佈。 並且,它的定價將維持與Fable 5完全相同(輸入 $10 / 輸出 $50 每百萬 Token)。

剛剛

Claude Code狂刪80%提示詞,Opus 5反手加回去了

Anthropic 在 Claude Code 中刪除了 80% 的提示詞,但隨著更強大的 Opus 5 模型推出,這些提示詞不僅被全數恢復,還新增了更多規範,導致「模型越強、規矩越多」的現象。業界分析認為,這反映出 Anthropic 在提升模型能力與維持使用行為可控性之間的反覆權衡,後續提示工程反而走向更嚴謹複雜的路徑。

剛剛