Open to workNew YorkGet in touch

Open-Weight Models

Kolibri model: 78B MoE, 3.46B active, Apache 2.0 weights for German and English

Aleph Alpha shipped a 78.1B-parameter, 3.46B-active German-English MoE under Apache 2.0. Here is what it costs to serve and when it beats Qwen3.6-35B-A3B.

Published
October 3, 2026
Read
6 min
Author
Samir Sengupta
Aleph Alpha's Kolibri German-English MoE model: 78.1B total, 3.46B active parameters, Apache 2.0 weights
Note 080 / 081Daily note · Written from 3 sources

the short version

  • Kolibri computes like a 3.5B model but needs the memory of a 78B one: about 78 GB of weights in FP8. GPU memory sets the hardware bill.
  • Its 128,000-token UniBPE tokenizer needs 11.2% fewer tokens for German than GPT-5's tokenizer, per Aleph Alpha. One independent test on the German constitution found about 15% fewer.
  • On vendor-run benchmarks it leads Qwen3.6-35B-A3B on math, GPQA and banking agents. It trails on LongBench Pro, AA-LCR, BFCL v4 and the AA-Omniscience Index.
  • It is the stronger pick for German-heavy, on-prem, regulated workloads. Qwen3.6-35B-A3B remains the lighter option for long-context retrieval and function calling.

Aleph Alpha released the Kolibri model on 3 October 2026. It is a German-English mixture-of-experts transformer with 78.1B total parameters and 3.46B active per token, with full weights on Hugging Face under Apache 2.0. It is trained natively to 262,144 tokens and was tested up to 1,048,576, and its weights take about 78 GB in FP8. On Aleph Alpha's own benchmarks it beats Qwen3.6-35B-A3B on AIME and GPQA in both languages, while trailing it on long-context and function-calling scores. It is worth choosing when the workload is mostly German, must run on your own hardware, and can afford the memory.

Aleph Alpha positions Kolibri for regulated work in public administration, industrials and aerospace. The model was built in Germany and trained on infrastructure in Germany and Finland. For teams whose workload is mostly German and must stay on their own hardware, there is now an Apache 2.0 option with a bilingual German/English tokenizer and a published account of how the training data was curated. The price is holding 78B parameters in memory to get 3B-class compute per token.

What the Kolibri model ships

  • Parameters: 78.1B total, 3.46B active per token (4.4%), across 50 layers that are all MoE, each with 384 experts, 6 routed per token, plus 1 shared expert.
  • Licence: Apache 2.0 for the weights and configuration files; Aleph Alpha keeps the rights to its training code and methods.
  • Context: 16,384-token pre-training context, longest trained length 262,144, tested up to 1,048,576.
  • Attention: sliding window of 512 tokens, with full attention every 5th layer (40 of 50 layers are sliding-window).
  • Tokenizer: 128,000-token vocabulary built with a new algorithm Aleph Alpha calls UniBPE.
  • Reasoning and tools: four effort levels (none, low, medium, high) and tool calling; knowledge cutoff 18 June 2026 for both languages.
  • Training data: 21.3% of pre-training tokens are German, with translation used for 6% overall; trained on 768 NVIDIA B200 GPUs.

The sources disagree on training tokens. The launch post says 20T pre-training tokens, distilled from over 200T tokens of raw data. The summary table in the tej.as write-up says about 24 trillion training tokens. The predecessor, Kolibri Origin, was a 30.6B-total, 3.27B-active model with a 65,536-token longest trained length and 7.51T pre-training tokens. It was never publicly released, and its pre-training finished on 11 June, three months before Kolibri's on 11 September.

Kolibri's tokenizer needs 11 to 15% fewer German tokens

Tokenizers trained mostly on English text chop up German compound words. GPT-5's o200k_base splits Bundesverfassungsgericht into 6 tokens, while Kolibri's tokenizer uses 2. Per the tej.as summary, UniBPE keeps the bottom-up merging of byte-pair encoding but scores each merge with the Unigram objective. Aleph Alpha's technical report says this needs 11.2% fewer tokens for German text than GPT-5's tokenizer, the best of the 9 others it measured.

The tej.as author ran an independent check on the full German Basic Law. Kolibri used 35,190 tokens, against 41,482 for o200k_base, 42,907 for Qwen3.5 35B-A3B, 43,478 for Mistral Small 4 and 43,850 for Gemma 4. That is about 15% fewer than GPT-5's tokenizer. On the official English translation it was effectively tied with o200k_base (39,875 vs 39,737). Fewer tokens means fewer decode steps for German output and more German per context window, but the test covers one legal document, and one Hacker News commenter asked whether the tokenizer is better in practice.

Kolibri leads on math, trails on long context

Aleph Alpha compares Kolibri with Qwen3.6-35B-A3B, Nemotron 3 Super 120B-A12B and Mistral Small 4 119B-A6B. Kolibri posts AIME 2025 96.9 (87.5 in German), GPQA diamond 84.3 (81.3 in German), LiveCodeBench v6 85.9 and τ³-bench banking 38.1, where the next best is Nemotron 3 Super at 15.5. Against Nemotron 3 Super, which has about four times its active parameters, Kolibri scores at or above it on every listed benchmark except HumanEval+ (92.7 vs 94.7).

Kolibri is behind Qwen3.6-35B-A3B on LongBench Pro (64.5 vs 70.8), AA-LCR (68.3 vs 69.7), BFCL v4 (61.4 vs 67.2), τ²-bench telecom (94.7 vs 99.1) and the AA-Omniscience Index (-32.8 vs -15.3). Aleph Alpha also reports internal customer-proxy suites. On the German public sector suite, scores rose from 0.54 to 0.75 across successive post-training runs. Those suites are internal, and every number in this section is vendor-run.

Memory, not compute, sets the hardware bill

The model card is explicit that the full model must be held in memory even though only part of it is active at any time. That comes to about 78 GB of weights in FP8. Aleph Alpha's serving argument is concurrency. On two H100s, a 123B variant it tested handled only 3 long-context 256k-token queries, while the 78B design handles 18 concurrent requests and decodes 28% faster. The four reasoning levels add a second lever, trading cost and latency against answer quality per request.

Neither the launch post nor the tej.as write-up, in the portions reviewed here, mentions official quantized releases, quality loss under quantization, or support in specific inference engines. That gap matters for anyone sizing hardware below the FP8 footprint.

What Hacker News commenters are pushing back on

One commenter on Hacker News questioned the intellectual-property-safety claim, arguing that a competitive model cannot be built without unlicensed training data. Other commenters disputed that. The sources partly speak to provenance. According to tej.as, the model card says English web text was rephrased with Gemma 4, German text with Mistral-NeMo, and Qwen3-32B labeled data for the quality filters, with the data then filtered for political bias. As that write-up puts it, sovereign does not mean nothing from outside Europe went in.

Another commenter judged Kolibri to perform about like Qwen3.6 35B A3B on the benchmarks. They said you would have to be a bit desperate to use it for coding, but suggested it would work for information extraction or classification in a regulated industry. They guessed it would run in roughly 96 GB of RAM with a decent quant, a figure no source verifies. A third commenter said Aleph Alpha recommended high-end setups without considering 4- or 8-bit quantization.

The sharpest practical complaint came from a user of a third-party hosted demo. They asked the model how to run itself on llama.cpp with less RAM than stated. With extended thinking on, they got little beyond advice to check the Hugging Face page. That is the flip side of the abstention training Aleph Alpha built with its Merlin-Arthur protocol, which teaches the model to say it does not know when the answer is not in the context. Neither source reports how often that abstention costs a correct answer.

When to choose it over Qwen3.6-35B-A3B

Kolibri makes sense when three conditions hold. Most of your input and output is German. The deployment has to stay on-premise under an Apache 2.0 licence, with a model built in Germany and trained in Germany and Finland. And you can hold about 78 GB of FP8 weights plus whatever your context lengths require. In that setting the tokenizer savings, the German AIME and GPQA scores, the banking and airline agent results, and the grounding-oriented abstention behaviour all favour it.

Qwen3.6-35B-A3B is the better default for long-document retrieval, function calling as measured by BFCL v4, and tighter memory budgets. Its name indicates about 35B total parameters, against Kolibri's 78.1B. Both comparisons rest on Aleph Alpha's own benchmark runs.

What is still unknown

Neither source reports independent benchmark reproductions, and the sector suites are internal. No source reviewed here mentions official quantized weights, quality-loss figures or support in particular inference engines. Quality between the 262,144-token trained length and the 1,048,576-token tested length is not documented in the reviewed material, and neither source gives a false-abstention rate for grounded question answering. The 20T versus 24T training-token figure is unresolved, so teams evaluating Kolibri should run their own German workloads at the quantization and context length they intend to serve.

Questions this raises

How much memory does the Kolibri model need?

The model card says the full model must be held in memory even though only 3.46B parameters are active per token. In FP8 that comes to about 78 GB of weights. The sources reviewed mention no official quantized releases.

Is Kolibri better than Qwen3.6-35B-A3B?

On Aleph Alpha's own benchmarks, Kolibri beats Qwen3.6-35B-A3B on AIME and GPQA in both English and German. It trails on LongBench Pro (64.5 vs 70.8), BFCL v4 (61.4 vs 67.2) and τ²-bench telecom (94.7 vs 99.1). All of these numbers are vendor-run.

What license is the Aleph Alpha Kolibri model under?

The weights and configuration files are released under Apache 2.0. Aleph Alpha keeps the rights to its training code and methods.

These daily notes are drafted by a model I run and operate myself - the same kind of pipeline this site is about - from sources published in the previous 24 hours, and every one lists what it read. The longer essays, the talks and the preprint are mine, written by hand.

More notes

Building something on this?

I ship production LLM, RAG and agentic systems for a living - the infrastructure behind the things these notes are about. Open to roles, contract work and research collaboration.