Kolibri is an open-weight German-English large language model from Aleph Alpha that uses a mixture-of-experts architecture: 78.1 billion parameters organized as 50 layers with 384 experts plus one shared expert per layer, where each token is routed to six experts, yielding about 3.46 billion active parameters per token. It was trained from scratch on ~24 trillion tokens (over 20% German) on infrastructure in Germany and Finland, ships under an Apache 2.0 weights license and occupies roughly 78 GB in FP8. Key engineering choices target German efficiency and long context: a 128k-token UniBPE tokenizer optimized for German compounds (measured to need ~11-15% fewer tokens on legal German than GPT-5’s tokenizer), a 262k-token native context window validated to 1,048,576 tokens through alternating sliding-window and occasional full-attention layers, tool-calling support, and a configurable reasoning_effort setting (none/low/medium/high).
Kolibri’s empirical strengths are strong German reasoning and math (top scores among similar active-size open models on AIME and German benchmarks), superior long-context performance on RULER at 1M tokens, and improved abstention behavior thanks to a Merlin‑Arthur training protocol that raises “I don’t know” responses (44% on a test) versus competitors. The sovereign positioning emphasizes EU-based build, legal compliance, and on-prem deployment, but sovereignty is qualified: some training pipelines used non‑European models for rephrasing and labeling. Trade-offs include the large in-memory footprint - full 78B weights must be hosted even though only part is active per token.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.