Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Aleph Alpha releases Kolibri for German and English AI

Kolibri brings open weights, reasoning and tool calling, with a million token context ceiling and practical deployment limits

Listen to this article

Aleph Alpha released Kolibri on October 3, making its new German and English reasoning model available to download. The launch announcement positions it for organizations that want to run AI on their own infrastructure, including public administration and industrial applications.

The release includes FP8 weights and a BF16 version under Apache 2.0. It supports adjustable reasoning effort and tool calling.

A large model with selective computation

The technical report describes a mixture of experts architecture with 78.1 billion total parameters and 3.46 billion active for each token. Each layer combines six selected experts with one shared expert. Forty of its 50 attention layers use a limited local window, while ten process the full context.

This design aims to limit the work required for each response. It still leaves a substantial model to store and serve, so the active parameter count alone is a poor guide to deployment requirements.

The context ceiling needs practical limits

The model card lists support for 1,048,576 tokens but recommends no more than 262,144 for serving efficiency and complex tasks. That lower figure is also its native trained context length. Teams should test representative documents before building around the maximum.

Aleph Alpha estimates a weight memory footprint of about 78 GB for FP8 and 156 GB for BF16. Those figures describe weights, leaving additional capacity needed for processing requests. This is a deployment planning issue even when only a fraction of the model is active.

What developers need to evaluate

The documented serving path uses Aleph Alpha’s inference package and its vLLM plugin. The guide explains reasoning and tool calling configuration, providing a path from downloaded weights to an application endpoint.

The company publishes performance comparisons using its own evaluation setup. Those results are useful starting points for testing, but they do not independently establish accuracy, reliability or costs in a customer’s deployment.

A practical trial should compare German and English output quality, document retrieval, tool selection and memory usage on the same tasks. Teams evaluating the wider ecosystem can also read our coverage of Ai2’s Olmo core 3 training tools.

Illustrative data center photograph by Brett Sayles on Pexels, used under the Pexels License. The photograph does not depict an Aleph Alpha facility.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile