Aleph Alpha has released Kolibri 1, a 78-billion-parameter mixture-of-experts model whose weights are available under the Apache 2.0 license. In a post on X, @Aleph__Alpha described it as a European-developed model with 3.46 billion active parameters and up to 1 million tokens of context. The supplied model pages position it for German- and English-language assistants, retrieval-augmented generation, coding, long-document processing and tool-calling workflows.

Video: Short launch video presents Kolibri 1’s parameter count, long context and Apache 2.0 release.

Load the post to view it as published on X. X may receive connection data.

View the original post on X ↗

What Aleph Alpha released

Kolibri 1 is available as an open-weight model. The Apache 2.0 license permits organisations to download and run the weights subject to that license’s terms, giving infrastructure teams more control over where inference takes place. The supplied model pages list the model’s deployment requirements and serving software, including Aleph Alpha’s inference package and a vLLM plugin.

The model is described as a mixture-of-experts (MoE) reasoning model. An MoE model contains multiple expert sub-networks and routes each token through only some of them. Kolibri 1 has 78,103,074,560 total parameters, while 3,457,573,120 parameters are active for each token.

That distinction is important for deployment. The active-parameter figure helps explain why the model can keep per-token computation relatively low compared with a dense model of the same total size. It does not mean that only 3.46 billion parameters must fit in memory: Aleph Alpha says the full model must be held in memory. The result is a trade-off between compute efficiency and a substantial memory requirement.

German and English capabilities

The model pages list German and English as Kolibri 1’s supported languages. Aleph Alpha describes that two-language focus as a deliberate choice to prioritise depth rather than broad multilingual coverage. The pages also describe a tokenizer tailored to German word structure and identify coding, structured extraction, retrieval-augmented generation and long-document processing among the intended tasks.

Kolibri 1 supports an explicit reasoning mode with configurable effort levels of low, medium and high. Thinking can also be disabled by setting the effort to none or passing enable_thinking=false. These settings control how much effort the model applies before producing its answer; they are not a guarantee that an output is correct.

The model also supports tool calling through an OpenAI-compatible API. A typical application could ask a German-language assistant about the current weather, allow the model to emit a structured get_weather call, execute that function, and send the result back for a final response. The calling system—not the model—must validate the function arguments and returned data before using them.

The 1-million-token context needs a qualification

Kolibri 1’s listed context length is 1,048,576 tokens. The same model pages recommend using contexts of 262,144 tokens or less for serving efficiency and complex tasks. Those figures describe different levels of support: Aleph Alpha says it validated quality and serving efficiency up to 1 million tokens, while the 262,144-token range is the recommended operating point for many deployments.

The pages say Kolibri 1 was pretrained on 16,384-token sequences, continued with 65,536-token sequences and then trained on 262,144-token sequences during a final long-context phase. Serving beyond 262,144 tokens requires additional configuration in the supplied vLLM instructions, including a larger maximum model length and a Hugging Face override for positional embeddings.

For an organisation, this means the headline 1-million-token limit should not be treated as a promise of uniformly efficient or suitable 1-million-token inference. The workload’s latency, throughput, prompt size and quality requirements still determine whether using that upper limit is practical.

Deployment requirements: FP8 versus BF16

Aleph Alpha supplies two weight distributions in the linked model pages. FP8 uses 8-bit floating-point weights, while BF16 uses the bfloat16 format. The pages present both as Kolibri 1 with 78 billion total parameters and 3.46 billion active parameters, but they list different memory footprints and GPU configurations.

Distribution

Weight format

Approximate model memory

Minimum hardware listed

Recommended hardware listed

Kolibri 1

FP8 weights with dynamically quantized activations and FP8 KV cache

About 78 GB

2× A100 80 GB; 2× H100 SXM5; 1× H200; 1× B200; or 1× B300

2× H100 SXM5; 2× H200; 1× B200; or 1× B300

Kolibri 1-BF16

BF16

About 156 GB

4× A100 80 GB; 4× H100 SXM5; 2× H200; 1× B200; or 1× B300

4× H100 SXM5; 2× H200; 2× B200; or 1× B300

The FP8 distribution therefore has the lower listed model-memory footprint, while the BF16 distribution requires roughly twice as much model memory according to the supplied pages. These figures describe the model weights and do not by themselves determine total system memory needs: context length, KV-cache settings, runtime overhead and workload concurrency also affect deployment.

The supplied serving instructions use Aleph Alpha’s aleph-alpha-inference package, which provides a Kolibri plugin for vLLM. For the FP8 distribution, the page gives this example for enabling reasoning and tool calling:

pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 \\
  --kv-cache-dtype fp8 \\
  --reasoning-parser kolibri1 \\
  --tool-call-parser kolibri1 \\
  --enable-auto-tool-choice

The BF16 model page gives a corresponding command using Aleph-Alpha/Kolibri-1-BF16 and does not include the FP8 KV-cache option. To serve contexts above 262,144 tokens, the supplied instructions add --max-model-len 1048576 and a max_position_embeddings override. Teams should test those settings against their own latency, memory and concurrency requirements rather than assuming that the maximum context is the best default.

The hardware lists include single-GPU options for some newer GPU models as well as multi-GPU configurations. In practical terms, the supplied configurations require one or more high-end data-centre GPUs, depending on the weight format and GPU model.

What Aleph Alpha reports in its evaluations

Aleph Alpha’s model pages include comparisons with other models across knowledge, mathematics, agentic tasks, coding, instruction following, grounding, retrieval and long-context evaluations. The FP8 model page reports Kolibri 1 post-training overall scores of 75.5 in English and 70.8 in German. The Kolibri-1-BF16 page reports 75.0 in English and 71.0 in German. The pages also report results for individual tests such as GPQA Diamond, AIME, HumanEval+, SWE-Bench Verified, retrieval tasks and long-context benchmarks.

These are company-reported measurements, not independent confirmation of general superiority. The supplied evaluation notes say that the comparisons used a shared evaluation setup for most benchmarks, while some tools used separate frameworks. They also explain that sampling parameters, context windows, reasoning effort and tool availability affect results, and that some models receive no score when they cannot call tools or when a prompt exceeds their context window.

The results are therefore most useful as evidence of the areas Aleph Alpha chose to measure under its documented setup. Organisations evaluating Kolibri should reproduce the tests that resemble their own German or English documents, retrieval corpus, coding tasks, tool schemas and latency targets before selecting it for production.

Aleph Alpha report graphic compares Kolibri with other models on throughput and evaluation scores.
Aleph Alpha report graphic compares Kolibri with other models on throughput and evaluation scores.

Image credit: @Aleph__Alpha on X

Intended use: human-reviewed systems

The supplied model pages describe Kolibri 1 as suitable for conversational assistants and agentic workflows in which a person reviews the model’s output before it is acted on. They also identify document-processing and drafting systems, question answering over an organisation’s own material, and internal knowledge and research tools as intended applications.

Tool calling can let a model request an API, search or code execution, but it does not make those actions safe or correct by itself. A production orchestration layer should validate tool arguments, constrain available operations, check returned data and decide whether a human must approve the next step. The model pages specifically place Kolibri on the advisory side of decision-support systems rather than making it the deciding component.

For deployment teams, Kolibri 1’s main proposition is therefore specific rather than universal: an Apache 2.0 open-weight model focused on German and English, with relatively low active computation per token, long-context support, reasoning controls and structured tool use. The corresponding cost is that the full 78-billion-parameter model still needs to fit in memory, with the supplied configurations requiring one or more high-end data-centre GPUs, depending on the weight format and GPU model. Its suitability depends on whether that infrastructure and a human-reviewed operating model match the organisation’s workload.

Sources