All posts
Open-WeightSelf-HostingRAGAI Architecture

Aleph Alpha Kolibri: German AI Model with Open Weights and Self-Hosting

6 min readThomas Stermole
Aleph Alpha Kolibri: German and English, open model weights under Apache 2.0 and deployment on your own infrastructure.

Kolibri is Aleph Alpha's new language model for German and English. Open weights and self-hosting make it interesting for companies that want to control their own AI processing. Selecting it also requires a look at quality in your actual workflow and the demands of operating it.

Aleph Alpha released Kolibri on 3 October 2026. This introduction puts the official specifications and model comparisons into context; it contains no hands-on test of my own. Sources checked on 5 October 2026. Official announcement

What is Aleph Alpha Kolibri?

Kolibri is a text-based mixture-of-experts model, or MoE. Only part of the model is activated for each generated token. This reduces computation compared with an architecture that uses all parameters. The complete weights still have to remain in memory. Model card

PropertyKolibri 1
LanguagesGerman and English
Total parametersAround 78 billion
Active parameters per tokenAround 3.46 billion
Context windowUp to 1,048,576 tokens; at most 262,144 recommended for efficient serving and complex tasks
CapabilitiesReasoning, tool calling, text processing
Model weights licenceApache 2.0

Source: Aleph Alpha's official model card.

The language focus gives companies a reason to shortlist Kolibri. An evaluation of your own must establish whether it reliably handles internal terminology, Austrian wording or incomplete documents. German-language support is a starting point, not proof of quality for every industry.

How does Kolibri compare with other models?

The charts below present a selection of Aleph Alpha's published comparison results. The values have been reproduced without changes and visualised anew for this article. These are vendor benchmarks, not independent measurements. Every axis starts at zero; higher values are better.

Aleph Alpha uses its own evaluation environments and, where applicable, each model's highest reasoning effort. The models differ in size and architecture. An overall score aggregates several tasks; it is not a success rate for your business process. Benchmark table and methodology

Overall scores in German and English

Vendor benchmark: Kolibri scores 70.8 overall in German and 75.5 in English; Qwen3.8 27B leads this selection with 79.9 and 80.2.Overall scores: a selection from Aleph Alpha's post-training benchmarks. Total and active parameters are shown for every model. Source: Aleph Alpha, 3 October 2026.

Kolibri scores above the MoE comparison models shown in this selection. The dense Qwen3.8 27B model achieves higher overall scores, but also activates substantially more parameters per token. This does not directly establish memory requirements or the price per request. Full comparison

My assessment: a model can be an attractive option without topping every ranking. What matters is whether the required quality can be achieved with acceptable latency, hardware and operational effort.

German industry RAG

With retrieval-augmented generation, or RAG, the model receives relevant document passages to ground its answer. This makes the comparison particularly interesting for internal knowledge systems.

Vendor benchmark for German industry RAG: Kolibri scores 67.5, Qwen3.5 70.0 and Qwen3.8 80.2; seven selected models are shown.Average across German Industry RAG tasks. These tasks are the vendor's internal customer-proxy benchmarks. Source: Kolibri Technical Report and model card.

Here, Qwen3.5 scores above Kolibri, even though Kolibri has the higher German overall score. The ranking therefore depends on the task. These industry tasks are tests developed by the vendor; they do not establish equivalent accuracy on your documents. Technical report, section 3.3.3

Document quality, retrieval and access permissions remain central to your own knowledge system. Changing the model does not fix missing sources or incorrect permissions. Further reading: RAG evaluation with a golden dataset.

Tool calling across multiple dialogue turns

Tool calling means that a model prepares calls to functions or APIs. For workflows, it matters whether this remains consistent across multiple steps.

BFCL v4 multi-turn vendor comparison: Kolibri scores 47.5; several models shown score higher, including Gemma 4 at 61.4 and Qwen3.5 at 59.9.BFCL v4, multi-turn: one tool-calling benchmark, not a general agent ranking. Source: Aleph Alpha's published benchmark table.

This comparison highlights a limitation: Kolibri does not lead the selection shown on this task. Teams planning a multi-step agent should examine precisely these cases. An impressive overall score is not enough to support the decision. Published tool-calling results

Self-hosting Kolibri: what does it involve?

Kolibri can run on your own infrastructure. Aleph Alpha provides a vLLM plugin and a documented inference environment. The server can be accessed through an OpenAI-compatible interface. Reasoning can be adjusted or disabled per request. Official inference repository

This gives companies control over model versions, access and the operating environment. It can reduce dependence on a single model endpoint. The company also assumes responsibility for capacity, updates, monitoring and availability.

Having 3.46 billion active parameters does not make Kolibri a small local model. The model card specifies around 78 GB for the FP8 weights alone and lists supported minimum configurations such as two A100 GPUs with 80 GB each or one H200. Runtime requirements, context and concurrent requests add to that. The published requirements cover the documented FP8 deployment; they do not confirm compatibility with arbitrary desktop LLM applications. Hardware specifications

Which companies should consider Kolibri?

I would investigate Kolibri when there is already a concrete German-language document or knowledge workflow and controlled self-hosting offers a clear benefit. Possible applications include internal research, document drafting and structuring information. These are potential use cases, not results I have measured.

For a modest start, first assess whether existing infrastructure and operating experience are sufficient. A large context window alone does not justify a major GPU purchase. If a smaller model already handles the task reliably, operating it may be more economical.

My next step would be a limited comparison using your own examples: complete documents, missing answers and error-prone tool calls. Assess output quality, response time and operating effort together. Kolibri belongs on the shortlist; vendor benchmarks do not justify replacing existing models across the board.

Frequently asked questions about Kolibri

What is Aleph Alpha Kolibri?

Kolibri is a language model from Aleph Alpha focused on German and English. It uses a mixture-of-experts architecture and supports reasoning and tool calling.

Can you self-host Kolibri?

Yes. The model weights are available under Apache 2.0. Aleph Alpha documents deployment through vLLM with its own plugin. The model card specifies around 78 GB of memory for the FP8 weights alone.

Is Kolibri better than Qwen, Gemma or Mistral?

That depends on the task. The published vendor benchmarks show strengths and weaknesses. They do not replace a comparison using your own documents, quality criteria and operational requirements.

Official sources

Next step

Sounds relevant for your company?

In a no-obligation initial call, we clarify within 30 minutes whether and where getting started is worthwhile for you — honestly and without sales pressure.

Request an initial call