Tejas Kumar examines Aleph Alpha’s Kolibri 1, an open-weight German-English large language model released on October 3, 2026 under the Apache 2.0 license for its weights and configuration files. The model has 78.1 billion total parameters but activates about 3.46 billion per token through a mixture-of-experts design; all weights still require roughly 78 GB of memory. It was trained from scratch on infrastructure in Germany and Finland with about 24 trillion tokens, including more than one-fifth German data, and has a native 262,144-token context window that Aleph Alpha validated up to 1,048,576 tokens. The article describes a 128,000-token UniBPE tokenizer designed for German compounds, which the author’s test found used 15% fewer tokens than GPT-5’s tokenizer on Germany’s Basic Law. Forty of 50 layers use 512-token sliding-window attention, while every fifth layer uses full attention to support long contexts. Kolibri can adjust reasoning effort across four levels, reasons in German on German prompts, and scored 87.5 on German AIME 2025 in Aleph Alpha’s evaluation. Aleph Alpha also trained it with the Merlin-Arthur protocol to acknowledge when supplied evidence does not support an answer; on Artificial Analysis’s Omniscience test, it abstained or partially answered 44% of questions it did not know, compared with 11.1% for Qwen3.5 35B-A3B and 23.7% for GPT-OSS 120B. The model led the company’s comparison on several German and English benchmarks, company-document questions, and one-million-token RULER performance, but performed worse on closed-book knowledge, multi-turn tool calling, coding-agent tests, and the 128,000-token RULER setting. It supports tool calling through an Aleph Alpha vLLM plugin and an OpenAI-compatible server, but requires data-center GPUs and had no hosted provider at launch. The article presents Kolibri as most suitable for German-language retrieval-augmented applications involving controlled, sensitive documents, while noting that its two-language scope, memory footprint, new serving stack, and weaker coding and general-knowledge results limit its use cases.
