Aleph Alpha Releases Kolibri-1: An Open-Weight Model With 78.1 Billion Parameters, Free on Hugging Face
The German company Aleph Alpha released Kolibri-1 on October 3, 2026, an open-weight model with weights published on Hugging Face under the Apache 2.0 license. All benchmark figures are, for now, the company's own, and no independent party has verified them.
The German AI company Aleph Alpha released Kolibri-1 on October 3, 2026 — an open-weight language model for English and German that the company positions as a "sovereign" alternative for public administration and regulated industries. The weights are on Hugging Face under the Apache 2.0 license, meaning organizations can download, run, and modify the model on their own hardware without asking permission. Alongside the model, the company published a 189-page technical report covering the architecture, training methods, and evaluation results.
The launch fell on German Unity Day, for a model aimed at public administration, industry, aerospace, and defense. But the timing is also unusual for another reason: Aleph Alpha is in the middle of a merger with the Canadian company Cohere, and this may be the company's last major independent release in its current form.
What Kolibri Is
Kolibri is a mixture-of-experts transformer (MoE) with 78.1 billion total parameters, of which only around 3.46 billion — roughly 4.4 percent — are active for each token. That is the point of the architecture: the model has the capacity of a large model but the inference costs of a much smaller one. Each of the 50 layers contains 384 routed experts, with six selected per token in addition to one shared expert. The attention mechanism is hybrid, with four sliding-window layers per full-attention layer, making long contexts cheaper to handle.
The German focus goes all the way down to the tokenizer. The company's own UniBPE tokenizer, with a vocabulary of 128,000 tokens, builds merges like BPE but evaluates each merge by Unigram loss. On German text it achieves 4.90 bytes per token, versus 4.35 for the GPT-5 tokenizer — that is, 11.2 percent fewer tokens on German web text. For a model meant to be used in German in the public sector, that means lower costs and faster processing on precisely the language that matters most.
Another detail stands out: the Merlin-Arthur protocol, a training method that teaches the model to abstain from answering when retrieved context does not support an answer. For users in government and other regulated environments, where fabricated answers can have consequences, it is a deliberate design decision.
The Training and the Sovereignty Argument
All technical figures below come from Aleph Alpha itself, relayed through secondary sources. The company states that the model was trained from scratch on a total of roughly 23.64 trillion tokens in three phases: the main phase lasted 21 days and covered 20 trillion tokens at a sequence length of 16,384, followed by 3.44 trillion tokens at a context of 65,536, and finally 201 billion tokens for long-context adaptation at 262,144.
The data mix is, according to the company, approximately 62.5 percent English, 23.9 percent German, and 13.6 percent code, plus over 2 trillion additional German tokens that are either curated from the web or synthetically generated. German thus makes up over 20 percent of the training data — unusually high for a model in this size class.
The sovereignty claim rests on more than language. The training ran on 768 Nvidia B200 GPUs on infrastructure in Germany and Finland, and Aleph Alpha states that the company controlled data curation, training, evaluation, and deployment end to end. The company is also a signatory of the EU's Code of Practice for general-purpose AI. It is worth noting that the GPUs are still American Nvidia cards — sovereignty here means control and location, not European hardware.
What It Costs to Run
For organizations considering running the model themselves, the requirements are moderate for the size class. The FP8 version takes up around 78 GB of memory and can be served on a single Nvidia H200, B200, or B300, or two A100 80 GB / H100 SXM5. The BF16 weights require around 156 GB. That means the model can run on a single modern accelerator or a pair of older GPUs — within reach of university hospitals, government administrations, and larger enterprises, not just hyperscalers.
The Benchmarks — With a Heavy Caveat
All results below are Aleph Alpha's own figures, published in the technical report and relayed via secondary sources. No independent benchmark lab has verified them, which matters more here than usual.
The reported numbers are strong on math and science: Aleph Alpha states 84.3 on GPQA Diamond and 96.9 on AIME 2025 in English. ForkLog, however, reports 84.8 on GPQA Diamond and 96.7 on AIME 2026, while MarkTechPost gives 96.0 on the same benchmark — the figures from different outlets thus diverge somewhat, and the reason is unclear. On the German-language version of AIME 2026, 90.2 is reported.
The weak points are just as telling. On agentic coding, Kolibri scores — according to the company's own figures — 27.7 on TerminalBench 2.1 versus 39.7 for Nemotron 3 Super, and 66.4 on SWE-Bench Verified versus 73.8 for Qwen3.6-35B-A3B. On BFCL v4 it scored 61.4 versus 70.5 for Qwen3.5. The dense Qwen3.8 27B scores higher overall in the comparison (80.2 EN, 79.9 DE) but activates around eight times more parameters per token — Kolibri's selling point, then, is efficiency, not absolute top performance. Trending Topics has, according to Startup Fortune, also pointed out that third-party comparisons place Kolibri behind today's leading open-weight models.
To Aleph Alpha's credit, the weak figures sit openly in its own report. But the picture is still one-sided: all positive and negative numbers come from the same source.
The Merger Shadow
The launch is happening as Aleph Alpha is in the process of being absorbed by Cohere. The companies disclosed merger plans in April 2026 and signed a definitive agreement on September 16. A person familiar with the matter has told The New York Times that the deal values the combined company at around $20 billion — a figure that rests on an anonymous source and should be treated as undocumented. The Schwarz Group, the German retail conglomerate behind Lidl and an existing Aleph Alpha investor, has according to several outlets injected $600 million as part of the deal.
OfficeChai describes the launch as happening "at an unusual time for the company, which is in the middle of being absorbed by Canada's Cohere." A very early Startup Fortune article stated that no parameter counts had been published at launch, most likely because the details came with the weights and the technical report shortly after.
The Open Questions
The most important caveat is verification: every technical specification and every benchmark result traces back to Aleph Alpha, via secondary coverage. Organizations considering Kolibri should read the 189-page technical report and the model card directly, and ideally run their own tests on their own workloads.
Then there is the strategic question: what happens to Kolibri under Cohere? None of the available sources say anything about a future roadmap after the merger — neither whether the Apache 2.0 releases will continue, nor whether the German sovereignty line will be maintained. For the German and European public-sector actors the model explicitly targets, that may be the most important unanswered question of all.

