Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters

2026年10月4日 07:01
站內 AI 整理稿

Aleph Alpha has released Kolibri, an open-weight Mixture-of-Experts (MoE) language model built for German and English.Kolibri has 78.1B total parameters but activates only 3.46B, or 4.4%, per token.

It accepts up to 1,048,576 tokens of context, lets users set reasoning effort per request, and ships under the Apache 2.0 license on Hugging Face.The target is sovereign deployment in regulated sectors such as public administration, industry and aerospace.Is it deployable?Yes.

The FP8 checkpoint is about 78GB and runs on a single B200, B300 or H200, or on 2 H100 SXM5 GPUs, served through vLLM with dedicated Kolibri reasoning and tool-call parsers.What is Kolibri?Kolibri (Kolibri-1) is a bilingual English-German MoE transformer developed end to end by teams in Germany.

According to the technical research report, Aleph Alpha team controlled the full pipeline: data, architecture, training infrastructure, post-training and evaluation.Training ran on infrastructure in Germany and Finland.

The design targets the EU General-Purpose AI Code of Practice, the EU AI Act and GDPR.Aleph Alpha is a signatory of that Code, and its data pipeline redacts personal data before training.

Architecture: Sparse Experts and Hybrid Attention Kolibri stacks 50 transformer blocks with a model width of 2,560.Every MoE layer scores all 384 routed experts with a sigmoid router, sends each token to the top 6, and always runs 1 shared expert.

Expert load is balanced with Exact Quantile Balancing and Load-Error Injection.Attention uses grouped-query attention with 48 query heads and 4 KV heads.Every fifth block uses full attention without positional encoding.

The other 40 blocks use sliding-window attention over the 512 preceding tokens, with RoPE.Sliding-window layers hold a fixed-size KV cache, so only 10 layers grow with context length.

At matched compute, Aleph Alpha team reports the hybrid supports sequences 4 times longer than a full-attention model.A Tokenizer Built for German The 128,000-token vocabulary is trained with UniBPE, which builds merges like BPE but scores each merge by Unigram loss.On German text it reaches 4.

90 bytes per token, versus 4.35 for the GPT-5 tokenizer.That means 11.2% fewer tokens on German web text.In English, Kolibri reaches 4.58 bytes per token against 4.67 for GPT-5.Training: 24T Tokens, Then SFT and RL Pre-training covered 20T tokens on 768 NVIDIA B200 GPUs, followed by 3.

44T mid-training tokens at 65,536 sequence length.A 201B-token long-context stage then trained on 262,144-token sequences.Aleph Alpha added more than 2T German tokens it curated from the web or generated synthetically.

Post-training combined supervised fine-tuning, mixed with MergeMix, with reinforcement learning on more than 1.2M internal tasks.The Merlin-Arthur protocol trains the model to abstain when retrieved context does not support an answer.Interactive Explainer: How Kolibri Works window.

addEventListener("message",function(e){if(e.data&&e.data.kolibriH){var f=document.getElementById("kolibri-xp");if(f)f.style.height=e.data.kolibriH+"px";}}); Benchmarks Aleph Alpha team evaluated every model with the same eval-framework setup.In English, Kolibri leads GPQA Diamond (84.

3), AIME 2025 (96.9) and AIME 2026 (96.0).It ties Qwen3.5 35B-A3B on the English agentic average at 63.4.It trails on BFCL v4, scoring 61.4 against 70.5 for Qwen3.5.The dense Qwen3.8 27B scores higher overall (80.2 EN, 79.9 DE), but activates about 8 times more parameters per token.

Against its internal predecessor Kolibri Origin, Kolibri decodes about 2.7 times more text per GPU while scoring 21.4 points higher in English.Kolibri vs Closest Competitors FeatureKolibri-1Qwen3.

6 35B-A3BNemotron 3 SuperMistral Small 4DeveloperAleph Alpha (Germany)Alibaba QwenNVIDIAMistral AI (France)Total / active params78.1B / 3.46B~35B / 3B120B / 12B119B / 6.

5BArchitectureMoE, sliding-window + full attentionMoE, Gated DeltaNet hybridMamba-2 + MoE + attentionMoEMax context1,048,576 tokens262,144 native, ~1M with YaRN1M tokens256k tokensReasoning controlnone / low / medium / highThinking on / offOn / off, low-effort modenone / highLicenseApache 2.

0Apache 2.0NVIDIA Nemotron Open Model LicenseApache 2.0Overall score (EN / DE)75.5 / 70.871.4 / 67.373.0 / 67.963.1 / 61.4Agentic avg (EN)63.462.154.940.7Industry RAG avg (DE)67.565.857.653.4Code avg (EN)89.387.788.382.0 Specs: model cards for Kolibri-1, Qwen3.

6-35B-A3B, Nemotron 3 Super and Mistral Small 4.

How to Deploy Kolibri Install the aleph-alpha-inference package, then serve with vLLM: Copy CodeCopiedUse a different Browserpip install 'aleph-alpha-inference>=1' vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \ --reasoning-parser kolibri1 --tool-call-parser kolibri1 \ --enable-auto-tool-choice The default context is 262,144 tokens.

Pass --max-model-len 1048576 with a maxpositionembeddings override for the full 1M window.Set reasoningeffort through chattemplate_kwargs.Aleph Alpha recommends temperature 1.0, top-p 0.97 and top-k 128.Key Takeaways 78.1B total, 3.

46B active: each token uses 6 of 384 routed experts plus 1 shared expert.40 sliding-window and 10 full-attention layers keep a 1M-token context affordable.Trained on 24T tokens, with German above 20% of the mix.Top Overall score in English (75.5) and German (70.

8) among the 12 MoE models Aleph Alpha compared.Apache 2.0 FP8 weights, serveable on a single B200 or H200 via vLLM.Check out the Model Weights and Technical Report.All credit goes to the researcher of this project.

Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?

Connect with us The post Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters appeared first on MarkTechPost.

Related

相關文章

鈦媒體生成式AI

當AI開始製造AI

15Bit2026.10.04 16:38 · 來自雲南全文3279字00:00 / 10:25當AI開始參與制造下一代AI,研發成果可能反過來加快研發,遞歸自我改進成為巨頭爭奪的新方向。但更快未必更好,誰定義進步,誰守住安全邊界?企業害怕落後的理由,也不能替代我們對這場加速是否值得的獨立判斷。文 | 15Bit劉慈欣在《鄉村教師》裡寫過一段對話。外星文明發現,人類沒有記憶遺傳,每一代都需要重新學習,卻依然發展出了能夠進入太空的文明。它們難以理解,這個物種怎樣積累知識。「他們沒有記憶遺傳,所有記憶都是後天取得的。

剛剛
IT之家生成式AI

System76 更新 COSMIC 項目 PR 模板,禁止貢獻者提交利用 AI 輔助完成的代碼

作者:漾仔 責編:漾仔 評論: 10 月 4 日消息,開發團隊 System76 現已更新 COSMIC 項目的 Pull Request(PR)模板,新增一份強制性檢查清單,明確禁止貢獻者提交利用 AI 輔助完成的代碼。貢獻者必須確認提交內容中不存在任何 AI 生成的代碼、註釋以及描述,否則無法按照要求提交貢獻。

剛剛
IT之家生成式AI

國內首部 AIGC 長劇《後西遊記》再上新,戰天宮篇今日開播

作者:浩渺 責編:浩渺 評論: 10 月 4 日消息,今日,國內首部 AIGC 長劇《後西遊記》再上新。10 月 4 日起,第一季 · 戰天宮篇將在湖南衛視、芒果 TV 雙平臺播出,10 月 9 日起,“劇好看”專區播出。注意到,神話題材季播劇《後西遊記》改編自明末清初匿名作者所著的同名神魔小說,其第一季共 30 集,每集 40 分鐘。

剛剛
鈦媒體生成式AI

殺豬盤用上了AI,騙局開始批量生產

豹變2026.10.04 11:19 · 來自四川全文5160字00:00 / 16:16殺豬盤換上AI引擎,普通人防不勝防。文 | 豹變,作者 | 高澤,編輯 | 邢昀在一個交友軟件上聊了幾天,對方始終回覆及時,話題也越聊越熟,甚至願意通一個視頻通話。如果你遇到這種場景,肯定不會懷疑對方是AI。但最近AI龍頭玩家Anthropic發佈的一份報告,揭開了另一種可能。

剛剛
鈦媒體生成式AI

李飛飛不等了

硅谷Tech news2026.10.04 08:38 · 來自北京全文5710字00:00 / 16:04“世界模型”被用濫了。大語言模型的紅利期進入後半段,物理世界成了巨頭們共同看中的下一個方向。資本入場比概念更早,僅今年一季度,全球物理AI領域的融資就超過64億美元,資金明顯向世界模型和基礎模型集中。讓AI在“腦子裡”預演物理世界、再到現實裡執行,被視作通向機器人、自動駕駛和下一代計算的必經之路。

3 小時前