# Aleph Alpha Kolibri is a 78.84 GB open-weight model, and its own documents disagree on German data URL: https://onticpost.com/aleph-alpha-kolibri-open-weights/ Published: 2026-10-05 Updated: 2026-10-05 By: Hill, Onticpost > Aleph Alpha Kolibri has 78.1 billion parameters in 78.84 GB of files, too big for 48 GB. We compared its card, announcement and report, and they disagree. Aleph Alpha Kolibri, a German and English language model with 78,103,074,560 parameters, was released on October 3, 2026, under the Apache 2.0 licence, according to the company’s model card. Aleph Alpha says 3.46 billion of those parameters are active for each token. The weights are 78.84 GB of files, and our test machine, an Apple M4 Pro, has 48 GB of memory. We have not run Kolibri, so this piece has no speed or memory figure of our own. Every claim about quality, training and hardware below is Aleph Alpha’s, taken from its model card, its announcement and its technical report. The file sizes and repository counts come from Hugging Face data we saved on October 5. **Takeaway points** - Hugging Face counts 78,103,074,560 parameters in the Kolibri-1 repository, and the model card says 3.46 billion of them are active for each token. - The model card sets the minimum hardware at one Nvidia H200 or two H100 cards, and serving needs Aleph Alpha’s own plugin for the vLLM engine. - We have not run Kolibri: its 78.84 GB of weight files is more than the 48 GB of memory on our test machine. ## What is Aleph Alpha Kolibri and how big is it? Kolibri is a mixture-of-experts model with 78.1 billion parameters in total, and the model card says it uses 3,457,573,120 of them per token. In a mixture-of-experts design each token is routed through a few small sub-networks instead of the whole network. The card lists 50 layers with 384 experts in each, of which 6 are chosen per token alongside 1 shared expert. The technical report puts the active share at 4.4%, and our own division of the two parameter counts gives 4.43%. The repository’s config file matches the card. It names the architecture Kolibri1ForCausalLM and gives a hidden size of 2,560, 48 attention heads, a vocabulary of 128,000 tokens and 262,144 maximum positions. Its list of layer types shows four sliding-window attention layers followed by one full attention layer, repeated ten times. The card says the sliding window covers 512 preceding tokens plus the current one, and the announcement says only 10 of the 50 layers process the full context. Aleph Alpha gives two context figures. The card calls 262,144 tokens the native context length, the length of the last training phase, and says the company has “validated quality and serving efficiency up to 1,048,576 tokens.” It recommends staying at or below 262,144 tokens for complex tasks and for deployments where speed matters. The model has a reasoning mode with four settings (none, low, medium and high) and supports tool calling, the card says. Its stated knowledge cutoff is June 18, 2026, in German and in English. The weights sit in two repositories, both under Apache 2.0 and neither gated. | Repository | Precision | Parameters | Size of weight files | | --- | --- | --- | --- | | Aleph-Alpha/Kolibri-1 | 8-bit floating point for most weights | 77,398,016,000 in 8-bit and 705,058,560 in 16-bit | 78.84 GB in 32 files | | Aleph-Alpha/Kolibri-1-BF16 | 16-bit | 78,103,074,560 | 156.21 GB in 32 files | Hugging Face records both repositories as created on October 2, a day before the release date on the card. ## How does Aleph Alpha Kolibri score against Qwen3.8 27B? Qwen3.8 27B scores higher than Kolibri overall in Aleph Alpha’s own table, 80.2 against 75.5 in English and 79.9 against 70.8 in German. The card says all models share one evaluation setup, with Kolibri set to high reasoning effort. We have not reproduced any score in this section. The card marks the best score only among the mixture-of-experts models and greys out the two dense ones, because “The dense models activate several times as many parameters per token”. Among the mixture-of-experts models Kolibri has the highest overall score in both languages. | Overall score | Kolibri | Qwen3.8 27B (dense) | Qwen3.5 35B-A3B | Nemotron 3 Super 120B-A12B | GPT-OSS 120B | | --- | --- | --- | --- | --- | --- | | English | 75.5 | 80.2 | 74.7 | 73.0 | 72.3 | | German | 70.8 | 79.9 | 69.8 | 67.9 | 70.2 | Aleph Alpha’s headline claim is about cost, not quality. Its announcement says “Kolibri sits on the Pareto frontier for quality versus serving cost, for both English and German.” It explains that none of the compared models delivers more quality at the same serving cost, or the same quality at lower cost. A reader who cares more about answer quality than serving cost should read the Qwen3.8 27B column too. Results vary a lot by task. The table below uses rows from the card, with the best score among the other mixture-of-experts models for comparison. | Benchmark | Kolibri | Qwen3.8 27B (dense) | Best other mixture-of-experts model | | --- | --- | --- | --- | | AIME 2025 (English) | 96.9 | 97.9 | 91.7 (Nemotron 3 Super) | | Tau3-Bench (Banking) | 38.1 | 50.0 | 16.0 (Gemma 4 26B-A4B) | | TerminalBench 2.1 | 27.7 | 76.8 | 39.7 (Qwen3.5 35B-A3B and Nemotron 3 Super) | | SWE-Bench Verified | 66.4 | 72.6 | 73.8 (Qwen3.6 35B-A3B) | | RGB Fact-Check (Error Correction) | 34.0 | 53.0 | 90.0 (Nemotron 3 Super) | Our reading is that Kolibri is strong at maths and at the banking tool-use test, and weak at terminal tasks and at correcting errors in documents it is given. Aleph Alpha’s two documents also disagree on one benchmark number. The card gives Qwen3.6 35B-A3B minus 12.5 on the AA-Omniscience Index, and the table in the announcement gives the same model minus 15.3. Kolibri’s own score is minus 32.8 in both. ## What hardware does Kolibri need, and will it fit a 48 GB Mac? No. The 8-bit weights are 78.84 GB of files, which is 30.84 GB more than the 48 GB of memory on our test machine, before the operating system and the context cache take anything. The card puts the memory footprint at about 78 GB. It lists the minimum hardware as two Nvidia A100 80 GB cards, two H100 SXM5 cards, or a single H200, B200 or B300, and recommends two H100 or two H200 cards or a single B200 or B300. Small active size does not shrink the file. The card states the trade-off itself: “the full model must be held in memory”. Our reading is that the 3.46 billion active parameters lower the compute for each token but do nothing for the memory a Mac has to find. For a dense model that does fit a 48 GB Mac, see [our test of Gemma 4 31B QAT on an M4 Pro with 48 GB](https://onticpost.com/gemma-4-31b-qat-m4-pro-48gb-tokens-per-second/). Serving is tied to one engine. The card says Kolibri needs Aleph Alpha’s own inference package, which provides a plugin for the vLLM engine, and it offers a container image or an install from the company’s code repository. It also gives recommended sampling settings of temperature 1.0, top-p 0.97 and top-k 128. ## Is there a GGUF or MLX build of Kolibri? Not from Aleph Alpha. The file lists of both official repositories hold safetensors files plus the config, tokenizer and licence, and no GGUF or MLX file. Those are the formats that llama.cpp, LM Studio, Ollama and Apple’s MLX read. Other accounts moved fast. A Hugging Face search on October 5 returned 32 repositories outside Aleph Alpha’s account with Kolibri-1 in the name. By name or tag, 12 are MLX builds, 5 are GGUF and 6 are NVFP4, a 4-bit format. The oldest appeared on October 3, the release day. We have not tested any of them, and Aleph Alpha’s card does not mention them. We cannot say whether a community build keeps the quality in the tables above. If you try a GGUF file, the steps in [how to import a local GGUF file into Ollama with a Modelfile](https://onticpost.com/import-gguf-ollama-modelfile-gemma-4/) apply, though that guide was written for Gemma 4 31B and not for this model. ## How was Kolibri trained, and why do Aleph Alpha’s documents disagree? Aleph Alpha says it trained Kolibri from scratch on about 24 trillion tokens, but its card, announcement and report give slightly different totals and two different German shares. The card says pre-training used 20 trillion tokens, mid-training 3.44 trillion and long-context training 201 billion. Those add up to 23.641 trillion by our arithmetic. The report states 24 trillion tokens as a flat figure, and the announcement says “nearly 24T tokens in total.” | Item | Model card | Announcement | Technical report | | --- | --- | --- | --- | | Total training tokens | 20T + 3.44T + 201B (23.641T by our sum) | “nearly 24T” | 24T | | Long-context tokens | 201B | 200B | 200B | | German share of pre-training tokens | about 23.9% | 21.3% | 21.3% | The card describes the pre-training mix as about 62.5% English, 23.9% German and 13.6% code. The announcement and the report both say German was 21.3% of pre-training tokens, and the announcement adds that this is about 4.3 trillion of the 20 trillion. We do not know which German figure is meant to stand. Part of the data is synthetic. The card says English web text was rephrased with Gemma 4 26B-A4B on 912 GPUs and German web text with Mistral NeMo 12B. The report says about 24% of the pre-training data, 4.8 trillion tokens, is synthetic in origin. On compute, the card reports pre-training on 768 Nvidia B200 GPUs for 21 days, or 392,000 GPU hours, then 90,000 for mid-training and 10,000 for the long-context stage. It estimates the energy at 950 MWh. That covers node power and data-centre overhead but leaves out fine-tuning, reinforcement learning and the smaller test models. The announcement says the model was trained on infrastructure in Germany and Finland. The card also records a limit of its data filtering. Its filters removed known adult sites and partially address child sexual abuse material, and “no dedicated CSAM detection was applied.” The training corpus was text only. Kolibri is the second model from this pipeline. Its predecessor, Kolibri Origin, has 30.6 billion parameters. The announcement says Origin finished pre-training on June 11 and Kolibri on September 11, and that Origin had no public release. ## Have we run Kolibri, and what has Aleph Alpha not published? We have not run Kolibri. We read three documents and two Hugging Face data files, and we compared file size with the memory of our test machine, an Apple M4 Pro with 48 GB. We took no speed or memory measurement, and we have not checked any benchmark score or training figure. Any speed figure we publish for a Kolibri build that fits would be measured the way described in [how to measure tokens per second in LM Studio without the cache fooling you](https://onticpost.com/measure-tokens-per-second-lm-studio/). Three things are missing from what Aleph Alpha has published. It has not released Kolibri Origin, so the gains the announcement reports over it cannot be checked outside the company. Its card gives no official build for llama.cpp or MLX and does not say whether one is planned. And its three documents do not reconcile the German data share, the token total or the AA-Omniscience figure for Qwen3.6 35B-A3B. For scale, [IQuest-Q1](https://onticpost.com/iquest-q1-open-weights/), another open-weight release we have reported on, comes to about 650 GB of files. ## Sources - [Aleph Alpha model card, Kolibri-1](https://huggingface.co/Aleph-Alpha/Kolibri-1) - [Aleph Alpha announcement, Kolibri Has Landed](https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/)