Aleph Alpha Releases Kolibri: A New Open-Weight MoE Model

At a Glance
| Item | Value |
|---|---|
| Published | 2026-10-03 |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code when the information was collected. Dates are JST.
Overview
German AI company Aleph Alpha has released “Kolibri," a new open-weight Mixture-of-Experts (MoE) model supporting English and German. This model consists of a total of 78.1 billion parameters with 3.46 billion active parameters per token, and supports a vast context length of up to 1 million tokens. Developed with European regulations and sovereignty in mind, its weights are provided on Hugging Face under the Apache 2.0 license.
Why It’s Buzzing
Kolibri is attracting attention because it combines a stance as a European-origin “sovereign AI" with high practicality in German-language environments. The development team emphasizes that the model was trained on infrastructure in Germany and Finland, mindful of European regulations such as the EU AI Act and GDPR, and without foreign control. Interest has also been drawn by the integration of practical technologies, such as the proprietary UniBPE tokenizer capable of efficiently processing long German compound words, and the implementation of the “Merlin-Arthur protocol," which trains the model to honestly answer “I don’t know" to prevent hallucinations. Furthermore, it features a function allowing users to select the depth of reasoning at inference time from four levels (none, low, medium, high), which has sparked interest for its ability to control the trade-off between cost and quality.
Discussion Points
Effectiveness as Sovereign AI and Corporate Background
Within the community, doubts have been raised regarding the sustainability of Aleph Alpha’s “sovereign" concept. Specifically, rumors suggest the company is in the process of merging with or being acquired by the Canadian company Cohere; pointing to this, critics question whether it can maintain its claim to European “sovereignty" if ownership largely shifts to Canada and operations are conducted from Toronto. Concerns and criticisms have also been raised regarding training data copyright and intellectual property safety, questioning whether others’ IP is being used for training without permission.
Performance Comparison with Competitor Models
There are also lukewarm evaluations of Kolibri’s benchmark scores. In official evaluations, it is noted that dense models like Qwen3.8 27B have outscored Kolibri in some tests. Additionally, complaints have been voiced that the selection of comparison targets may be biased, as official reports lack comparisons with recent lightweight MoE models like Qwen3.8 Flash, instead comparing it against models from about a year ago.
Hardware Requirements and MoE Trade-offs
Active discussions are also taking place regarding the hardware requirements to run Kolibri. While it is very lightweight and capable of fast processing with 3.5 billion active parameters during token processing, due to MoE characteristics, all 78 billion parameters of the entire model must be held in GPU memory. Consequently, running it requires about 78 GB of VRAM, making it difficult to operate on personal Macs or laptops. The community has discussed the desire for 4-bit or 8-bit quantized versions for easier testing, alongside methods to suppress VRAM requirements.
Evaluation of Proprietary Technology and Open Source
On the other hand, many favorable opinions have been expressed regarding its technical approach and documentation publication stance. In particular, the 189-page technical report is detailed enough to feel like a tutorial on “how to make a modern agentic LLM," and has been highly praised for its rare openness, covering even the dataset creation methods. Interest has also focused on the ingenuity of the UniBPE algorithm—which can represent text in German contexts with significantly fewer tokens than conventional tokenizers—and the design of the “Merlin-Arthur protocol" as a countermeasure against hallucinations.
Community Reactions
Evaluation of Sovereign AI’s Effectiveness and Corporate Background
Regarding the stance of championing European sovereignty, both lukewarm skepticism and sympathy can be seen. Mentioning the rumored acquisition or merger by the Canadian company Cohere, opinions question whether the grounds for calling it “sovereign" will waver if the operational structure becomes North American-led. Conversely, some showed a positive reception, understanding the need for independent AI foundations in countries outside the US and China, and welcoming cooperation between Germany and Canada, as well as with other European companies like Mistral.
Concerns regarding intellectual property were also discussed. Pointing out that in the competitive landscape of 2026, building a competitive model without widely collecting others’ data is extremely difficult, severe views emerged questioning whether data rights issues have truly been resolved even with public sector involvement. Furthermore, regarding the description that existing LLMs (Large Language Models) were used for paraphrasing training data, voices questioned transparency, asking which models were used and whether that might cause the model to learn specific models’ expression styles.
Pros and Cons of Performance Comparisons with Competitor Models
On the performance front, severe criticisms have emerged regarding comparisons with competing open models. Highlighting that in the company’s benchmark environment, the German evaluation score of the dense model Qwen3.8 27B (79.9) surpassed Kolibri (70.8), some voices poked fun at the results failing to match other countries’ models despite appealing to “sovereignty." Dissatisfaction was also expressed that the comparison targets were nearly year-old older models, lacking comparisons with the latest Qwen3.8 Flash which has a close active parameter count.
At the same time, some opinions evaluate the model’s characteristics from a practical standpoint. Analyses suggest that while it may lag behind other models in memory (closed-knowledge tests), coding performance, and multi-turn tool calling, it could find sufficient practical business value in use cases such as providing external documents for fact-checking or acting in roles like a “trust adapter" to monitor and audit outputs of other models. Cautious views were also expressed regarding its proprietary evaluation axis of Pareto optimality, seeking to determine whether it is a sharp achievement specialized in a specific domain or merely local optimization.
Hardware Requirements and MoE Practicality
Regarding hardware requirements, perplexity at the cumbersomeness specific to the MoE architecture stands out. Even with a small inference active parameter count of 3.46 billion (approx. 3.5 billion), roughly 78 GB of VRAM must be secured to load the entire set of weights, and regret was expressed that individual developers cannot easily run it on local laptops or Macs. Multiple requests were made for official low-bit quantized versions, such as 4-bit or 8-bit, since the provided weights center around fp16 to curb VRAM consumption.
On the other hand, positive reports regarding inference speed came from developers who actually ran it on data center-grade hardware. Reports indicated that generation speeds of approximately 170 tokens per second were achieved at FP8 precision in an RTX 6000 Ada generation environment. However, a caution shared was that in that test environment, the model tended to consume excessive tokens on the thought process, causing reasoning to become overly prolonged.
Praise for Technical Report Transparency and Proprietary Technology
Great praise has poured in from the community regarding the technical publishing stance. Regarding the 189-page technical report, it details even the dataset construction procedures, leading to expressions of awe that it reads like a DIY tutorial for modern agentic LLMs. Researchers and engineers highly praised it, noting that a report openly disclosing the development process to this extent is unprecedented.
Positive reactions were also widely directed at the technical ingenuity. The UniBPE tokenizer, which considers German grammatical structures, and the “Merlin-Arthur protocol," which increases the probability of answering “I don’t know" when a context lacks grounding, were received as extremely clever approaches to counter hallucinations in RAG (Retrieval-Augmented Generation). Moreover, proactive moves to actually test the model were seen, such as volunteer developers launching free hosting environments for the community to try it without a GPU.
What to Read Next
- How to read Perplexity → Benchmark glossary

