Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloads
Reflection AI has introduced Beam, its first open-weight model.Beam is a sparse Mixture-of-Experts (MoE) model with 501B total parameters and 23B active per token, built for coding, reasoning and agentic workloads.
As per the Reflection AI team, Beam directly competes with larger open models like GLM 5.2 while using 3 to 4x less inference compute on reasoning benchmarks.Is it deployable today?Not for self-hosting yet.Beam is in final red-teaming.Early access runs through a waitlist on the Reflection platform.
What is Reflection Beam?Beam is a general agent model trained from scratch by Reflection AI.It targets enterprise coding and agentic workloads.Reflection positions Beam as advancing the Western open-weight frontier.The research team is candid about the gap.
Kimi K3 stays ahead on raw capability, so Beam’s pitch is efficiency at inference time.Users get a reasoning effort parameter.Lower settings favor short answers.Higher settings allow longer reasoning on hard tasks.Teams can match effort to task difficulty and compute budget.
How Beam was Pretrained Beam was pretrained on 23.8 trillion tokens from the web, public sources and proprietary licensed datasets.Reflection team states that its curation removed about 95% of raw internet tokens.It also kept roughly 1.
8 trillion high-quality tokens that conventional filters would have dropped.The architecture interleaves local and global attention with fine-grained routed experts.Load balancing builds on auxiliary-loss-free balancing from DeepSeek-V3, adding cosine decay of expert-bias updates.
The busiest expert reached just 1.04x average load at the end of pretraining.Across all 52 layers, residual norms stayed bounded using depth-based scaling, SandwichNorm, attention gating and FP32 residual accumulation.Pretraining finished in under 4 weeks on 6,144 NVIDIA GB300 NVL72 GPUs.
Goodput reached 92.3% near the end, with 9 semi-automatic rewinds.Midtraining extended effective context to 1M tokens.High-Compute Reinforcement Learning RL is Beam’s central scaling axis.The run used 10.5K NVIDIA GB300 GPUs for 4 weeks and generated over 100 million rollouts.
Maximum rollout context was 256K tokens.Training and grading consumed about 1.3 billion sandboxes across nearly 1 million coding, agentic and STEM environments.Reflection trained with fully asynchronous policy gradients.Every token is tagged with the policy version that produced it.
New algorithms kept learning stable even at one-day staleness, 107 weight versions behind the current policy.The team reports no plateau as RL compute increased.Infrastructure numbers are notable.The system sustained 110K concurrent rollouts on average.
New weights reached the inference fleet in a median of about 12 seconds.71 inference incidents were handled without stopping training.A controllable length penalty taught Beam to solve tasks with fewer tokens.
Browsing skills also improved without browsing tasks in the RL mix, which suggests transfer across agentic domains.Safety and Alignment Reflection trained a separate safety and alignment teacher from the pretrained checkpoint.
It merged that teacher with the RL teacher using multi-teacher on-policy distillation.Safety training used deliberative alignment.Safety evaluation results will appear in the technical report.Benchmarks (Reflection-Reported) On SWE-bench Verified, Beam scores 80.9 versus 70.7 for Nemotron 3 Ultra.
On Terminal Bench v2.1, Beam scores 80.1, close to GLM 5.2 at 81.0.DeepSeek V4.1 Flash (90.6) and Kimi K3 (88.3) lead there.These numbers come from Reflection’s table, which sources rival scores from Artificial Analysis and DataCurve.
Beam vs Closest Open-Weight Competitors FeatureReflection BeamGLM-5.2Nemotron 3 UltraDeepSeek V4.1 FlashKimi K3DeveloperReflection AI (US)Z.ai (China)NVIDIA (US)DeepSeek (China)Moonshot AI (China)Total params501B~753B550B552B backbone + 196B Engram2.
8TActive params23B~40B55B8B prefill / 16B decode104BContext1M (effective)1MUp to 1M1M1MInputTextTextTextText + imageText + imageLicenseApache 2.0 (planned)MITOpenMDW-1.1MITKimi K3 LicenseWeightsLater in Oct 2026AvailableAvailableAvailableAvailableTerminal Bench v2.180.181.056.490.688.
3SourceReflectionHugging FaceNVIDIAHugging FaceHugging Face Scores as published in Reflection’s Beam announcement.Specs verified October 5, 2026.Beam is the smallest model here by total parameters.Its 23B active count sits below GLM-5.2, Nemotron 3 Ultra and Kimi K3.Apache 2.
0 and MIT are standard permissive licenses.Kimi K3’s custom license adds attribution requirements for very large products.window.addEventListener('message',function(e){if(e.data&&e.data.mtpBeamH){var f=document.getElementById('mtp-beam-frame');if(f)f.style.height=e.data.
mtpBeamH+'px';}}); Key Takeaways Beam is a 501B MoE with 23B active parameters, text-only, and 1M effective context.Pretraining used 23.8T tokens on 6,144 NVIDIA GB300 GPUs in under 4 weeks.RL ran 100M+ rollouts on 10.5K GB300 GPUs over 4 weeks.Reflection reports 80.9 on SWE-bench Verified and 80.
1 on Terminal Bench v2.1.Apache 2.0 weights are scheduled for later this month.Check out the Technical details and Early Access.The post Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloads appeared first on MarkTechPost.
Related
相關文章

最火AI崗位FDE:月薪5萬,都幹這些…
FDE(前線部署工程師)是近期最受關注的AI職位之一,海外年薪中位數約20萬美元,國內大廠也開出月薪三到五萬元。這份工作強調駐場梳理客戶的業務本體(Ontology),溝通時間佔七成以上,開發僅約三成。從業者認為,FDE與傳統外包不同,關鍵在於能否將經驗沉澱回自家產品並複用。
DeepSeek Harness v0.2 為其開源代理框架帶來官方桌面應用程式
DeepSeek 已為 DeepSeek Harness (dsh) 推出官方桌面應用程式,dsh 是其開源代理框架。該應用程式隨 v0.2 預覽版一同發布,安裝檔支援 macOS(Apple 晶片)與 Windows(64 位元)。目前可作為預覽版部署使用,使用者可從 deepseek.com/harness 下載,或執行 npx @deepseek-ai/dsh web。DeepSeek 提醒未來可能會有破壞相容性的變更。 v0.2 新增內容:框架是將模型轉化為代理的執行環境,能讀取檔案、執行指令並維持計畫。v0.2 預覽版針對日常工作與程式開發進行優化。內建功能包括:預載常用辦公室與開發工具;新增插件管理頁面,可安裝、設定、啟用與停用插件;以及右側邊欄提供檔案與差異審查預覽。

DeepSeek擴招!彈性計算團隊大量HC,尤其需要資深工程師
彈性計算團隊大量HC招人!尤其需要資深工程師。三週前不是剛招過一輪嗎?咋又缺人了。這次沒發崗位JD,直接甩了一篇DeepSeek技術分享: 《DeepSeek彈性計算(DSec):面向大規模Agent訓練的沙盒基礎設施》 現在DeepSeek的一套DSec擴展分片,大約有160臺服務器、3萬個CPU核心和250TB內存,每天要服務約300萬個沙盒。

OpenAI安全團隊持續地震!負責人離職,三名員工因洩密被開
OpenAI安全團隊再傳人事動盪,負責安全透明度的David Robinson已離職,OpenAI尚未公布繼任者。Robinson團隊主要負責撰寫模型「風險說明書」system card,對外解釋內部安全評估與部署決策。此外,另有三名員工因洩密遭到開除,顯示團隊內部問題持續延燒。

蘋果 MacBook Pro 外接 iPhone 17 Pro Max 運行 AI 模型,預填充性能最高提升 44%
作者:遠洋 責編:遠洋 評論: 10 月 4 日消息,Qwen3.8‑27B 是一款性能不錯的人工智能模型,但前提是設備擁有足夠的顯存與系統內存。搭載 M4 Pro 芯片的 MacBook Pro 僅有 24GB 統一內存,在運行該 AI 模型時,預填充速度會因此受到限制。
IBM Brings Bob to Self-Hosted and Air-Gapped Environments: Agentic Software Development Without Moving Your Code
IBM has made a self-hosted deployment option for IBM Bob generally available. Bob is IBM’s agentic software development platform. It covers the full lifecycle: understanding code, planning work, executing changes and validating results.