Aleph Alpha releases Kolibri for German and English AI
Kolibri brings open weights, reasoning and tool calling, with a million token context ceiling and practical deployment limits

Aleph Alpha released Kolibri on October 3, making its new German and English reasoning model available to download. The launch announcement positions it for organizations that want to run AI on their own infrastructure, including public administration and industrial applications.
The release includes FP8 weights and a BF16 version under Apache 2.0. It supports adjustable reasoning effort and tool calling.
A large model with selective computation
The technical report describes a mixture of experts architecture with 78.1 billion total parameters and 3.46 billion active for each token. Each layer combines six selected experts with one shared expert. Forty of its 50 attention layers use a limited local window, while ten process the full context.
This design aims to limit the work required for each response. It still leaves a substantial model to store and serve, so the active parameter count alone is a poor guide to deployment requirements.
The context ceiling needs practical limits
The model card lists support for 1,048,576 tokens but recommends no more than 262,144 for serving efficiency and complex tasks. That lower figure is also its native trained context length. Teams should test representative documents before building around the maximum.
Aleph Alpha estimates a weight memory footprint of about 78 GB for FP8 and 156 GB for BF16. Those figures describe weights, leaving additional capacity needed for processing requests. This is a deployment planning issue even when only a fraction of the model is active.
What developers need to evaluate
The documented serving path uses Aleph Alpha’s inference package and its vLLM plugin. The guide explains reasoning and tool calling configuration, providing a path from downloaded weights to an application endpoint.
The company publishes performance comparisons using its own evaluation setup. Those results are useful starting points for testing, but they do not independently establish accuracy, reliability or costs in a customer’s deployment.
A practical trial should compare German and English output quality, document retrieval, tool selection and memory usage on the same tasks. Teams evaluating the wider ecosystem can also read our coverage of Ai2’s Olmo core 3 training tools.
Illustrative data center photograph by Brett Sayles on Pexels, used under the Pexels License. The photograph does not depict an Aleph Alpha facility.



