Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters
Aleph Alpha has released Kolibri, an open-weight Mixture-of-Experts (MoE) language model built for German and English.Kolibri has 78.1B total parameters but activates only 3.46B, or 4.4%, per token.
It accepts up to 1,048,576 tokens of context, lets users set reasoning effort per request, and ships under the Apache 2.0 license on Hugging Face.The target is sovereign deployment in regulated sectors such as public administration, industry and aerospace.Is it deployable?Yes.
The FP8 checkpoint is about 78GB and runs on a single B200, B300 or H200, or on 2 H100 SXM5 GPUs, served through vLLM with dedicated Kolibri reasoning and tool-call parsers.What is Kolibri?Kolibri (Kolibri-1) is a bilingual English-German MoE transformer developed end to end by teams in Germany.
According to the technical research report, Aleph Alpha team controlled the full pipeline: data, architecture, training infrastructure, post-training and evaluation.Training ran on infrastructure in Germany and Finland.
The design targets the EU General-Purpose AI Code of Practice, the EU AI Act and GDPR.Aleph Alpha is a signatory of that Code, and its data pipeline redacts personal data before training.
Architecture: Sparse Experts and Hybrid Attention Kolibri stacks 50 transformer blocks with a model width of 2,560.Every MoE layer scores all 384 routed experts with a sigmoid router, sends each token to the top 6, and always runs 1 shared expert.
Expert load is balanced with Exact Quantile Balancing and Load-Error Injection.Attention uses grouped-query attention with 48 query heads and 4 KV heads.Every fifth block uses full attention without positional encoding.
The other 40 blocks use sliding-window attention over the 512 preceding tokens, with RoPE.Sliding-window layers hold a fixed-size KV cache, so only 10 layers grow with context length.
At matched compute, Aleph Alpha team reports the hybrid supports sequences 4 times longer than a full-attention model.A Tokenizer Built for German The 128,000-token vocabulary is trained with UniBPE, which builds merges like BPE but scores each merge by Unigram loss.On German text it reaches 4.
90 bytes per token, versus 4.35 for the GPT-5 tokenizer.That means 11.2% fewer tokens on German web text.In English, Kolibri reaches 4.58 bytes per token against 4.67 for GPT-5.Training: 24T Tokens, Then SFT and RL Pre-training covered 20T tokens on 768 NVIDIA B200 GPUs, followed by 3.
44T mid-training tokens at 65,536 sequence length.A 201B-token long-context stage then trained on 262,144-token sequences.Aleph Alpha added more than 2T German tokens it curated from the web or generated synthetically.
Post-training combined supervised fine-tuning, mixed with MergeMix, with reinforcement learning on more than 1.2M internal tasks.The Merlin-Arthur protocol trains the model to abstain when retrieved context does not support an answer.Interactive Explainer: How Kolibri Works window.
addEventListener("message",function(e){if(e.data&&e.data.kolibriH){var f=document.getElementById("kolibri-xp");if(f)f.style.height=e.data.kolibriH+"px";}}); Benchmarks Aleph Alpha team evaluated every model with the same eval-framework setup.In English, Kolibri leads GPQA Diamond (84.
3), AIME 2025 (96.9) and AIME 2026 (96.0).It ties Qwen3.5 35B-A3B on the English agentic average at 63.4.It trails on BFCL v4, scoring 61.4 against 70.5 for Qwen3.5.The dense Qwen3.8 27B scores higher overall (80.2 EN, 79.9 DE), but activates about 8 times more parameters per token.
Against its internal predecessor Kolibri Origin, Kolibri decodes about 2.7 times more text per GPU while scoring 21.4 points higher in English.Kolibri vs Closest Competitors FeatureKolibri-1Qwen3.
6 35B-A3BNemotron 3 SuperMistral Small 4DeveloperAleph Alpha (Germany)Alibaba QwenNVIDIAMistral AI (France)Total / active params78.1B / 3.46B~35B / 3B120B / 12B119B / 6.
5BArchitectureMoE, sliding-window + full attentionMoE, Gated DeltaNet hybridMamba-2 + MoE + attentionMoEMax context1,048,576 tokens262,144 native, ~1M with YaRN1M tokens256k tokensReasoning controlnone / low / medium / highThinking on / offOn / off, low-effort modenone / highLicenseApache 2.
0Apache 2.0NVIDIA Nemotron Open Model LicenseApache 2.0Overall score (EN / DE)75.5 / 70.871.4 / 67.373.0 / 67.963.1 / 61.4Agentic avg (EN)63.462.154.940.7Industry RAG avg (DE)67.565.857.653.4Code avg (EN)89.387.788.382.0 Specs: model cards for Kolibri-1, Qwen3.
6-35B-A3B, Nemotron 3 Super and Mistral Small 4.
How to Deploy Kolibri Install the aleph-alpha-inference package, then serve with vLLM: Copy CodeCopiedUse a different Browserpip install 'aleph-alpha-inference>=1' vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \ --reasoning-parser kolibri1 --tool-call-parser kolibri1 \ --enable-auto-tool-choice The default context is 262,144 tokens.
Pass --max-model-len 1048576 with a maxpositionembeddings override for the full 1M window.Set reasoningeffort through chattemplate_kwargs.Aleph Alpha recommends temperature 1.0, top-p 0.97 and top-k 128.Key Takeaways 78.1B total, 3.
46B active: each token uses 6 of 384 routed experts plus 1 shared expert.40 sliding-window and 10 full-attention layers keep a 1M-token context affordable.Trained on 24T tokens, with German above 20% of the mix.Top Overall score in English (75.5) and German (70.
8) among the 12 MoE models Aleph Alpha compared.Apache 2.0 FP8 weights, serveable on a single B200 or H200 via vLLM.Check out the Model Weights and Technical Report.All credit goes to the researcher of this project.
Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?
Connect with us The post Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters appeared first on MarkTechPost.
Related
相關文章

System76 更新 COSMIC 項目 PR 模板,禁止貢獻者提交利用 AI 輔助完成的代碼
作者:漾仔 責編:漾仔 評論: 10 月 4 日消息,開發團隊 System76 現已更新 COSMIC 項目的 Pull Request(PR)模板,新增一份強制性檢查清單,明確禁止貢獻者提交利用 AI 輔助完成的代碼。貢獻者必須確認提交內容中不存在任何 AI 生成的代碼、註釋以及描述,否則無法按照要求提交貢獻。

國內首部 AIGC 長劇《後西遊記》再上新,戰天宮篇今日開播
作者:浩渺 責編:浩渺 評論: 10 月 4 日消息,今日,國內首部 AIGC 長劇《後西遊記》再上新。10 月 4 日起,第一季 · 戰天宮篇將在湖南衛視、芒果 TV 雙平臺播出,10 月 9 日起,“劇好看”專區播出。注意到,神話題材季播劇《後西遊記》改編自明末清初匿名作者所著的同名神魔小說,其第一季共 30 集,每集 40 分鐘。

殺豬盤用上了AI,騙局開始批量生產
豹變2026.10.04 11:19 · 來自四川全文5160字00:00 / 16:16殺豬盤換上AI引擎,普通人防不勝防。文 | 豹變,作者 | 高澤,編輯 | 邢昀在一個交友軟件上聊了幾天,對方始終回覆及時,話題也越聊越熟,甚至願意通一個視頻通話。如果你遇到這種場景,肯定不會懷疑對方是AI。但最近AI龍頭玩家Anthropic發佈的一份報告,揭開了另一種可能。

GPT-6要“吃掉”3D公司?這家公司不到2年ARR翻百倍,破1億美元
通用AI模型GPT-6 Astra能直接生成3D場景,但與專門的AI 3D生成工具Meshy相比,在角色與複雜道具的細節刻画上仍有明顯落差。儘管通用模型能力快速進展,Meshy的業務反而持續成長,不到兩年時間ARR從100萬美元翻百倍突破1億美元。

