Editorial illustration for Aleph Alpha's New 78B-Parameter Model Supports 4x Longer Sequences
Aleph Alpha's 78B Model Handles 4x Longer Sequences
Aleph Alpha put out Kolibri this week, an open-weight Mixture-of-Experts model built specifically for German and English, and the numbers are the story. The model carries 78.1 billion parameters total but only switches on 3.46 billion of them, about 4.4%, for any given token. Context window runs up to 1,048,576 tokens, four times the stretch of what most comparable open models handle, and users can dial reasoning effort up or down per request. It's released under Apache 2.0 on Hugging Face.
The FP8 checkpoint weighs in around 78GB and fits on a single B200, B300 or H200, or spreads across two H100 SXM5 GPUs, served through vLLM with parsers built for Kolibri's own reasoning and tool-call formats. Aleph Alpha, a German company, built the whole pipeline in-house, data collection through evaluation, and trained the model on infrastructure in Germany and Finland. That's a deliberate choice aimed at sovereign deployment for public administration, industry and aerospace clients who need to stay inside EU rules. Aleph Alpha has signed onto the EU's General-Purpose AI Code of Practice, and the training data goes through redaction before it ever touches the model.
Aleph Alpha has released Kolibri, an open-weight Mixture-of-Experts (MoE) language model built for German and English. Kolibri has 78.1B total parameters but activates only 3.46B, or 4.4%, per token. It accepts up to 1,048,576 tokens of context, lets users set reasoning effort per request, and ships under the Apache 2.0 license on Hugging Face.
Why this matters
Kolibri is a test of whether "sovereign AI" can mean something beyond a procurement slide. Aleph Alpha built this for German public administration and aerospace contracts, not for leaderboard bragging rights, and the engineering choices back that up: 3.46B active parameters out of 78.1B keeps inference costs down, the million-token context window matters for processing long regulatory or technical documents, and the UniBPE tokenizer tuned for German suggests someone actually cared about efficiency on a language English-first labs tend to treat as an afterthought. The Apache 2.0 license and Hugging Face release mean developers can actually run this, not just read a paper about it.
For founders building in regulated European markets, this is worth testing directly rather than taking Aleph Alpha's 4x-longer-sequence claim at face value. Compute-matched comparisons are easy to frame favorably. Researchers should pay attention to the MoE routing efficiency at 4.4% activation and whether that reasoning-effort toggle holds up under real workloads. The real question isn't the parameter count, it's whether a German-focused sovereign model finds customers willing to trade raw capability for data residency and licensing control.
Common Questions Answered
What is the Mixture-of-Experts architecture used in Aleph Alpha's Kolibri model?
Kolibri is a Mixture-of-Experts model with 78.1 billion total parameters but only activates 3.46 billion parameters, or 4.4%, for any given token. This selective activation approach significantly reduces inference costs while maintaining model capability, making it more efficient than traditional dense models.
How does Kolibri's context window compare to other open-weight models?
Kolibri supports a context window of up to 1,048,576 tokens, which is four times longer than what most comparable open-weight models can handle. This extended context window is particularly valuable for processing long regulatory documents and technical materials used in German public administration and aerospace applications.
What languages is Aleph Alpha's Kolibri model optimized for?
Kolibri is built specifically for German and English, with a UniBPE tokenizer tuned for German to ensure optimal performance in both languages. This language-specific optimization reflects Aleph Alpha's focus on serving German public administration and aerospace contracts rather than pursuing generic leaderboard performance.
Can users adjust the reasoning effort when using Kolibri?
Yes, Kolibri allows users to dial reasoning effort up or down on a per-request basis, providing flexibility in balancing computational resources with response quality. This feature enables users to optimize performance based on their specific needs and available computational capacity.
Under what license is Kolibri released and where can it be accessed?
Kolibri is released under the Apache 2.0 open-source license and is available on Hugging Face as an open-weight model. This makes it freely accessible to developers and organizations who want to use or modify the model for their applications.
Further Reading
- Kolibri: A Sovereign European Model on the Pareto Frontier - Aleph Alpha
- Aleph Alpha puts Kolibri's full weights on Hugging Face under Apache 2.0 - StartupFortune
- Aleph Alpha releases Kolibri, a 78B-parameter open-weight AI model built in Europe - CryptoBriefing
- Aleph Alpha Releases Kolibri for Self-Hosted German and English AI - SuperpowerDaily
- Aleph Alpha Releases Kolibri-1, an Open German Reasoning Model With 1M-Token Context - AlphaSignal