Aleph Alpha launches Kolibri-1 as an open-weight German-English reasoning model designed for self-hosted enterprise and public-sector deployments. Released on October 3, the model combines 78.1 billion total parameters with just 3.46 billion active parameters per token and supports contexts up to one million tokens.

 

The Heidelberg company published downloadable FP8 and BF16 weights on Hugging Face, an extensive model card and deployment instructions for vLLM. Kolibri-1 arrives as European governments and regulated industries seek models they can operate on infrastructure they control rather than depend entirely on hosted US services.

 

The release has four defining specifications:

  • 78.1 billion total parameters in a mixture-of-experts architecture.
  • 3.46 billion active parameters for each generated token.
  • A 262,144-token native context validated up to 1,048,576 tokens.
  • Apache-2.0 weights for German and English deployments.

 

Aleph Alpha Launches Kolibri-1 as Open Weights

Kolibri-1 is available for local serving rather than only through a vendor-controlled application programming interface. Developers can run it through Aleph Alpha's inference package and a compatible vLLM server, with an OpenAI-compatible endpoint for existing applications.

 

The model supports explicit reasoning settings, including low, medium and high effort, plus a mode that returns an immediate answer without a visible reasoning phase. Its Hermes-style tool calling can produce structured requests for external functions, searches, code execution or business systems.

 

Aleph Alpha recommends human review before Kolibri-1's outputs trigger real actions. The model card positions it for document processing, retrieval-augmented generation, internal research, drafting and agent workflows where an application validates tool results, not for unsupervised decision-making.

 

The Apache-2.0 grant is narrower than a fully open training stack. It covers the published weights and configuration files, while Aleph Alpha retains rights to other code, architecture details, parameter settings and training methods that are not included in the repository.

 

How Kolibri-1 Activates 3.46 Billion Parameters

Kolibri-1 uses a 50-layer mixture-of-experts transformer. Each expert layer contains 384 routed experts, but the router selects six of them for a token alongside one shared expert. That sparse design reduces computation during inference while keeping the representational capacity of a much larger model.

 

The trade-off is memory. All 78 billion parameters still need to be available even though only a fraction is active at once. Aleph Alpha lists an FP8 memory footprint of roughly 78 GB and recommends at least two 80 GB A100 GPUs, two H100s, one H200, one B200 or one B300.

 

Most attention layers use a 512-token sliding window, while every fifth layer can attend across the full context. This 4:1 balance limits the cost of processing long documents without completely removing global attention.

 

The training corpus contained about 20 trillion pre-training tokens, followed by 3.44 trillion tokens in mid-training and 201 billion for the long-context stage. Aleph Alpha says the main pre-training run used 768 Nvidia B200 GPUs for 21 days, excluding the later phases.

 

Kolibri-1 Extends a 262K Native Window to One Million Tokens

The model was trained first on 16,384-token sequences, then at 65,536 tokens and finally at a native maximum of 262,144. Aleph Alpha validated quality and serving efficiency at 1,048,576 tokens without applying the usual position-scaling technique.

 

That one-million-token headline needs context. The company recommends staying at or below 262,144 tokens for complex tasks and latency-sensitive production workloads. The larger window is useful for unusually long archives or codebases, but it demands more memory and does not guarantee that every detail will be recalled equally well.

 

Aleph Alpha's RULER evaluation reports a score of 69.8 at 256,000 tokens, 65.5 at 512,000 and 63.2 at one million. The declining scores show that maximum input length and reliable long-context performance are related but distinct measures.

 

German specialization is another deliberate constraint. Roughly 23.9% of the pre-training corpus was German, 62.5% English and 13.6% code, and the tokenizer was built to handle German word structure efficiently. The company chose depth in two languages rather than broad multilingual coverage.

 

Benchmarks Frame Kolibri-1 for Sovereign AI Deployments

In Aleph Alpha's evaluation harness, Kolibri-1 reached an overall score of 75.5 in English and 70.8 in German. Those results exceeded Qwen3.6 35B-A3B in the same comparison, but remained below Qwen3.8 27B, which scored 80.2 and 79.9 respectively.

 

The model scored 96.9 on the English AIME 2025 mathematics benchmark and 94.7 on Tau2-Bench Telecom tool use. It was less dominant on TerminalBench 2.1, scoring 27.7 against 76.8 for Qwen3.8 and 39.7 for both Qwen3.5 and Nemotron 3 Super.

 

These are developer-reported results rather than independent certification. Aleph Alpha used the same harnesses and settings across compared models where possible, but missing scores, different supported context windows and reasoning configurations still complicate direct rankings.

 

The strategic pitch is therefore broader than benchmark leadership. Kolibri-1 gives European administrations, manufacturers and regulated businesses a downloadable bilingual model with a documented training summary, local deployment path and EU-oriented compliance disclosures.

 

Its timing also matters. Aleph Alpha is pursuing a merger with Canadian model developer Cohere, creating a transatlantic enterprise AI company with a Heidelberg research center. Kolibri-1 demonstrates that the German operation is still producing foundation-model research while that transaction awaits regulatory clearance.

 

Adoption will depend on whether outside teams reproduce the claimed efficiency and quality on real German-language workflows. The immediate evidence will come from Hugging Face downloads, community deployment reports and evaluations conducted beyond Aleph Alpha's own benchmark suite.

 

More on This Story